Google DeepMind seeks a Staff Software Engineer for Inference Performance Optimization to accelerate large scale AI inference in Mountain View. You will analyze the full inference stack, design optimization techniques, and build metrics to raise throughput per GPU and reduce latency. Requirements include a technical bachelor’s degree or equivalent, 8+ years of software development, strong Python and C++ skills with experience in serving codebases, and knowledge of AI model execution constraints and modern serving architectures. Preferred candidates have real world LLM inference experience, profiling and observability expertise, and familiarity with accelerator concepts such as memory bandwidth and quantization. To apply, quantify impact with metrics and tailor your resume to show measurable improvements while aligning with safety and ethics.
At Google DeepMind our mission is to build the world's first general-purpose learning agent. Central to this mission is the complex task of measuring the intelligence of our prototypes. As a Software Engineer, you will be working with the cutting edge AI agents developed by our exceptional team of Machine Learning and Neuroscience research scientists. Your responsibilities will include everything from creating systems for agent testing using 2D and 3D games to developing test problems within physics simulators. You will create graphical visualization of results, build competitive agent leaderboards and test new algorithms on robots. To succeed in this role you will need to have a strong foundation in software engineering and enjoy working on a wide range of challenging problems within a mission-driven team.
As an Inference Performance Engineer, you will push the boundaries of AI model execution at scale. In this role, you will be at the forefront of making large-scale AI inference faster, cheaper, and more efficient. You will analyze the entire inference stack to identify critical bottlenecks and drive systemic improvements. By combining deep systems profiling, benchmarking, and first-principles problem solving, your work will directly maximize hardware throughput, reduce cost-to-serve, and empower our cross-functional teams to make data-driven capacity and latency tradeoffs.
Artificial intelligence will be one of humanity’s most transformative inventions. At Google DeepMind, we are a pioneering AI lab with exceptional interdisciplinary teams focused on advancing AI development to solve complex global challenges and accelerate high-quality product innovation for billions of users. We use our technologies for widespread public benefit and scientific discovery, ensuring safety and ethics are always our highest priority.