Google is seeking a Staff Software Engineer for Data Acquisition Infrastructure in San Jose, onsite, full time. The role demands deep technical chops with 8 years of C++ programming, 5 years building large scale infrastructure or distributed systems, and hands on experience with API design, infra spanner, distributed processing and Generative AI agents, plus familiarity with server platform or pod. Preferred qualifications include a Master’s degree or PhD, strong data structures and algorithms, and proven ability to lead cross team efforts delivering high availability services with a focus on production excellence. You’ll help design and scale crawling, rendering and scraping infrastructure powering Search, model training and more. To apply, tailor your resume to show scalable systems impact, leadership, reliability and concrete metrics, highlight cross team collaboration, and be ready for system design, coding and leadership questions.
Google's software engineers develop the next-generation technologies that change how billions of users connect, explore, and interact with information and one another. Our products need to handle information at massive scale, and extend well beyond web search. We're looking for engineers who bring fresh ideas from all areas, including information retrieval, distributed computing, large-scale system design, networking and data storage, security, artificial intelligence, natural language processing, UI design and mobile; the list goes on and is growing every day. As a software engineer, you will work on a specific project critical to Google’s needs with opportunities to switch teams and projects as you and our fast-paced business grow and evolve. We need our engineers to be versatile, display leadership qualities and be enthusiastic to take on new problems across the full-stack as we continue to push technology forward.
In this role, you will be working on the large scale of crawling, rendering, scraping infrastructure for Google, which acquires tens of billions of Web content per day, to power Google products including Search, Gemini model training, Google Cloud, Ads, Shopping, etc.
We are seeing increasing needs from many teams within Google to acquire various content from the Web to power their products (e.g., news pages, web automation (agent), model training data, etc.). Our web discovery and crawling infrastructure is positioned to support these emerging cases, and we are investing in building scalable and sustainable web navigation functionalities.