- Company: Get A Job.ai
- Location: London
- Salary: Pay not listed
Website Get A Job.ai
Represented by Get A Job.ai
Responsibilities
Our talent team is representing a confidential AI research organization seeking a Staff Infrastructure Engineer to lead cluster infrastructure strategy in London. You will own the technical direction for agent-driven cluster lifecycle management across multiple cloud providers and private datacenters.
- Define and execute the technical roadmap for automated cluster provisioning, updates, and decommissioning at massive scale
- Partner with cross-functional teams to ensure new compute capacity is delivered on schedule
- Design and implement high-bandwidth inter-cluster connectivity solutions leveraging cloud platforms
- Collaborate with security teams to ensure all clusters are provisioned with secure-by-default configurations
- Drive strategy on cluster scalability, fault tolerance, and operational homogeneity
- Work closely with cloud providers and internal research, inference, and product teams to shape long-term infrastructure strategy
- Establish operational excellence practices including incident response, postmortem culture, and on-call health
- Mentor and coach engineers through technical leadership and hands-on guidance
What we're looking for
Required qualifications:
- Deep expertise in distributed systems, reliability engineering, and cloud platforms (Kubernetes, infrastructure-as-code, AWS/GCP/Azure)
- Strong proficiency in at least one systems language (Rust, Go, or Python) plus Terraform experience
- Proven track record leading complex, multi-quarter technical initiatives spanning multiple teams
- Ability to build alignment across senior stakeholders and communicate effectively at all organizational levels
Preferred qualifications:
- 10+ years of software engineering experience, including time as a technical lead setting team direction
- Experience operating large-scale compute infrastructure at hyperscale (100+ clusters, 10K+ nodes)
- Depth in Kubernetes internals, cluster provisioning and management systems, or cluster orchestration systems
- Experience with cloud networking: VPC design and peering, Shared VPC/Transit Gateway, Cloud Interconnect/Direct Connect, Cloud NAT, cross-cloud private connectivity, BGP and route control, edge load balancing and DDoS mitigation
- Experience with cluster and host networking: CNI (Cilium), eBPF, NetworkPolicy, multi-NIC, sFlow, service mesh (Istio/Envoy/Linkerd, mTLS)
- Experience with cluster security: pod security standards and admission control, RBAC and least-privilege IAM, node and container hardening, supply-chain/image provenance
- Deep experience with infrastructure-as-code (Terraform, Atlantis) and workflow orchestration (Temporal, Argo Workflows)
- Skill in quickly understanding systems design tradeoffs and tracking rapidly evolving software systems
Education and logistics:
- Bachelor's degree or equivalent combination of education, training, and experience in a relevant field
- Location-based hybrid policy: expected in-office presence at least 25% of the time in London
- Visa sponsorship available for qualified candidates
How we work with you
Candidates apply directly through Get A Job.ai. Our recruiting team will screen qualified applicants and coordinate interviews with our client. We handle all scheduling and feedback communication throughout the process. Please do not attempt to contact the client directly, as all applications must be submitted through our platform to be considered.
Pay
This role offers competitive compensation. Specific pay details will be discussed with qualified candidates during the screening process based on experience and qualifications.
Equal Opportunity: Get A Job.ai is committed to providing equal employment opportunities to all qualified applicants regardless of race, color, religion, sex, national origin, age, disability, or any other protected characteristic. We encourage applications from candidates of all backgrounds.
Apply with Get A Job.ai
A recruiter will review your profile and submit you to the client. Do not contact the client directly.
Apply through Get A Job.ai. A recruiter will review your profile and submit you. Do not contact the client directly.
Apply through Get A Job.ai. A recruiter will review your profile and submit you.
Local insights for this role are preparing — this section updates automatically in a few seconds (or refresh).
Listing facts
- Role Staff Infrastructure Engineer, Cluster Infrastructure
- Employer Get A Job.ai
- Location London
- Type Full Time
- Pay (from listing) Pay not listed
- Posted September 16, 2026
- Apply by October 16, 2026
- Overview Full job description on this page (494 words)
Facts above come from this job record on Get A Job.AI — not copied from third-party review sites.
Limited public data for this employer
We only show facts we can ground in public sources (Wikidata, O*NET, news/discussion links, or this listing). We do not invent Glassdoor-style ratings, salaries, or testimonials when data is thin. Use the listing facts, occupation context, and related openings below while we continue researching.
Explore related openings
Keep exploring on Get A Job.ai
Not quite the right fit? Your next opportunity is a click away.
- Browse all jobs
- More jobs by category
- Remote jobs you can do from anywhere
- Research typical pay for this role
- Set a job alert so new matches reach you first
- Upload your resume to apply faster
Hiring instead? Post a job and reach candidates searching right now.
