Software Engineer (Compute Efficiency), London
Isomorphic LabsIsomorphic Labs is looking for a Software Engineer (Compute Efficiency) to join the Compute Infrastructure team in London. The core mission is to keep a planet-scale accelerator fleet operating at peak health and efficiency, pairing with in-house ML platforms, the performance and scaling squad, and AI researchers to squeeze more throughput from every GPU-hour. The company is applying frontier AI to speed up drug discovery and, ultimately, solve disease, with a culture that values interdisciplinary collaboration and meaningful impact.
IsoLabs, launched in 2021, builds on and beyond the Nobel-winning AlphaFold system to accelerate scientific discovery across multiple modalities. Its world-leading drug design engine uses AI models to design novel molecules, anticipate drug performance, and drive rational medicines for serious diseases. The team pursues digital biology to shorten the path from research to real-world therapies, with a clear mission to one day solve all disease with AI.
The role sits within a culture that blends curiosity, collaboration, and ambitious problem-solving. It’s a place for engineers who want to shape the infrastructure that makes rapid, responsible AI-driven science possible.
Keeping the accelerator fleet humming
In this position, the engineer designs and operates observability systems and telemetry pipelines to monitor compute efficiency, hardware health, and workload throughput across distributed clusters. The work combines building reliable monitoring with ongoing efforts to keep the fleet healthy as new hardware and architectures come online, so performance scales alongside research needs.
The job also centers on driving hardware efficiency and node reliability across the accelerator fleet, collaborating with ML platforms and infrastructure teams to reduce waste and push utilization higher. It’s about identifying bottlenecks and partnering with researchers and platform engineers to improve the throughput of both training and inference workloads, so researchers can move faster without sacrificing stability.
Day-to-day duties include maintaining and hardening cloud infrastructure and cluster deployments used for research, development, and production, and working with science, research, product, business development, and operations teams to inform decisions about tooling, infrastructure, and architecture. The role also involves contributing to core technical choices that guide how the team approaches instrumentation, tooling, and system design.
What you’ll bring to the team
The essential capabilities for this role include:
- Hands-on experience running, monitoring, and debugging infrastructure for large-scale AI and ML workloads
- Background in cloud compute infrastructure design, with a preference for Google Cloud Platform
- Strong programming skills across languages commonly used for infrastructure and tooling
- Extensive experience deploying and operating Kubernetes at scale
- Familiarity with Nvidia GPU generations
- Proven track record building production observability and telemetry stacks
Nice-to-haves
- Experience as an ML software engineer or infrastructure site reliability engineer on related projects
- Conceptual understanding of ML workload efficiency paradigms
- Experience leading projects and delivering outcomes to multidisciplinary stakeholders
- Familiarity with Google TPU generations
- Knowledge of workload scheduling, ML-driven R&D cycles, ML efficiency research, and hardware benchmarking
Hybrid work in London
The company values in-person collaboration and follows a hybrid model that requires being in the London office three days a week, Tuesdays, Wednesdays, and one other day based on your team. The role is London-based and onsite, with flexibility to discuss accommodations if needed during the screening process. When you submit an application, your data will be processed in line with the privacy policy. privacy policy.
Advice for applicants
Lead with a resume line that shows you’ve built or run production-grade AI/ML infrastructure at scale, and make sure Kubernetes at scale and cloud design, especially on GCP, are front and center.
Back up must-have skills with concrete results. Include metrics like improvements in GPU utilization, reductions in downtime, or increases in training or inference throughput tied to your work.
Be ready to walk through a specific instance where you identified compute waste, traced it to a bottleneck, and implemented a fix that measurably boosted efficiency or throughput.
Ask about how the team defines and tracks workload goodput, and what plans exist to expand the accelerator fleet or adopt new hardware. A good starter question: how do you measure the impact of observability improvements on daily research cycles?