Staff Engineer, Compute Foundation
StripeStripe is hiring for a Staff Engineer to join its Compute Foundation in Sydney, Australia, with on-site work. Stripe builds the payments infrastructure that millions of businesses rely on to accept payments and grow revenue. Its mission is to increase the GDP of the internet, and this role sits at the core of that effort. The Compute Foundation team is responsible for the systems that let Stripe engineers ship code, configurations, and infrastructure changes safely and quickly. Within that group, the Imaging team designs, builds, and distributes Linux machine images, manages host bootstrapping and OS configuration, maintains OS packages and security patching, and provides platforms that let service teams run workloads across cloud providers and regions. If you enjoy shaping the ground floor of a large, service-driven platform, this could be a great fit.
The role centers on owning the foundation of Stripe’s compute fleet. You’ll influence the base operating system, the package ecosystem, security posture, and bootstrap behavior that Stripe’s services depend on. You’ll lead architectural transformations such as modernizing the OS configuration layer, containerizing workloads, and scaling machine-image distribution across multi-cloud regions while tightening OS upgrades and patching. You’ll manage the broad, end-to-end host lifecycle, from image construction and package distribution to kernel and OS configuration, early boot with systemd, storage, networking initialization, and workload readiness, tackling tradeoffs around reliability, security, compliance, developer experience, and performance. Your judgment will help contain risk and prevent fleetwide incidents by designing validation, canarying, rollout, health evaluation, and rollback mechanisms. As a Staff engineer, you’ll set technical direction for the Imaging team, author designs that span multiple teams, and mentor engineers who rely on your technical north-star.
Laying the groundwork for Stripe’s compute
You’ll own the foundation of Stripe’s compute fleet: the host infrastructure that boots reliably, receives accurate configurations, and operates safely. Your decisions shape the base operating system, the package ecosystem, security posture, and bootstrap behavior across Stripe’s fleet.
You’ll lead active architectural transformations: the team is modernizing the OS configuration layer, containerizing workloads, scaling machine image distribution across multi-cloud regions, and streamlining OS upgrades and critical security patching.
You’ll drive broad, end-to-end host lifecycle ownership: work across image construction, package distribution, kernel and OS configuration, early boot/systemd, storage, networking initialization, and workload readiness. You’ll tackle complex engineering challenges that balance reliability, security, compliance, developer experience, and performance.
Your judgment prevents fleetwide incidents: changes to machine images, packages, kernels, bootstrap systems, and host configuration can have broad consequences. You will design validation, canarying, rollout, health-evaluation, and rollback mechanisms that contain blast radius and make foundational infrastructure changes safer.
Agency to shape technical strategy: as a Staff engineer, you will set technical direction for the team’s systems, author designs that span multiple teams, mentor engineers, and be the person engineering managers and engineers rely on for the technical north-star.
Foundations you’ll need
Must-have essentials for this role include a lengthy track record in professional software engineering with a focus on production infrastructure systems at scale, and a proven ability to lead large, ambiguous projects from design to delivery while coordinating cross-team migrations. You should have deep Linux expertise, covering boot, initialization, systems, processes, networking, filesystems, storage, packages, and production debugging, with strong experience in cloud compute infrastructure, including machine images, virtual machines, autoscaling fleets, launch configurations, block storage, IAM, health evaluation, and multi-region operation in AWS, Azure, or an equivalent environment. Experience designing or operating fleet-scale image, host-management, configuration-management, or infrastructure-provisioning systems is essential, as is familiarity with safe rollout practices such as automated validation, canarying, progressive rollout, health evaluation, rollback, and blast-radius containment. A solid background in service reliability and operational excellence, leading incident response, reducing toil, and building reliable, debuggable, maintainable systems, is required. Finally, you should demonstrate a track record of broad technical impact across multiple large systems, with fluency across a complex codebase, a strong ability to mentor through code reviews, and capability to steer technical direction beyond just execution.
- Over ten years of professional software engineering experience, with a proven history of designing and delivering production infrastructure at scale
- Demonstrated success leading large, ambiguous infrastructure initiatives from design through to delivery, including coordinating migrations across many consuming teams
- Deep Linux expertise across boot processes, system initialization, networking, storage, and production debugging, plus solid cloud compute know-how for building and operating machine images across multi-region setups in AWS, Azure, or similar
- Background designing or operating fleet-scale image, host-management, configuration-management, or infrastructure-provisioning systems
- Experience with safe rollout approaches, including automated validation, canarying, progressive rollout, health checks, rollback, and blast-radius containment
- Proven ability to lead incident response and reduce toil, delivering reliable, debuggable, and maintainable systems
- Ability to influence and set technical direction across multiple teams through code reviews, mentorship, and architectural leadership
Bonus skills that help
- Experience building modern image-based, immutable, or declarative infrastructure platforms
- Hands-on with containers, Kubernetes, and moving VM workloads to managed platforms
- Experience operating large Linux fleets across AWS, Azure, or multiple regions
- Knowledge of progressive delivery, automated validation, fleet health signals, and rollback systems
- Background modernizing legacy configuration-management or host-provisioning systems
- Understanding OS security patching, software supply-chain security, and artifact provenance
- Strong developer-platform instincts, with experience creating safe abstractions and self-service workflows
Logistics at a glance
The role is based on-site in Sydney, Australia.
Advice for applicants
Lead your resume with a project that shows end-to-end ownership of host infrastructure, from image construction and distribution to kernel and OS configuration and rollout. Emphasize any work that involved multi-region or multi-cloud deployments and the impact you had on reliability and patching processes.
When you describe your must-have skills, call out specific Linux areas you’ve owned, like boot sequences, systemd startup, networking initialization, and how you validated changes before broad rollout. Include concrete examples of machine-image decisions and how you balanced security with developer experience.
Prepare to discuss architecture and strategy in interviews, focusing on how you would set the direction for Stripe’s host-infrastructure stack, what guardrails you’d put in place for safe changes, and how you would measure fleet health and rollback readiness.
Consider asking about the Imaging team’s current roadmap, how they measure blast radius, and how they collaborate with security and compliance teams to maintain a strong security posture across regions.