Loading...

Staff Software Engineer, AI Reliability Engineering

  • Company: Get A Job.ai
  • Location: London
  • Salary: Pay not listed
  • Full Time
  • London

Website Get A Job.ai

Represented by Get A Job.ai

Responsibilities

Our client, a leading AI technology organization, is seeking a Staff Software Engineer to join their AI Reliability Engineering team in London. In this role, you will:

  • Develop appropriate Service Level Objectives for large language model serving systems, balancing availability and latency with development velocity
  • Design and implement monitoring and observability systems across the token path
  • Assist in the design and implementation of high-availability serving infrastructure across multiple regions and cloud providers
  • Lead incident response for critical AI services, ensuring rapid recovery, thorough incident reviews, and systematic improvements
  • Support the reliability of safeguard model serving, critical for both site reliability and the organization's safety commitments

What We're Looking For

We are representing a confidential AI organization in their search for candidates who have:

  • Strong distributed systems, infrastructure, or reliability backgrounds—we're looking for reliability-minded software engineers and SREs
  • Curiosity and courage to jump into unfamiliar systems during an incident and help drive resolution even without deep expertise yet
  • Holistic thinking about how systems compose and where the seams are
  • Ability to build lasting relationships across teams—the engagement model depends on being welcomed as teammates, not outsiders with opinions
  • Care about users and feel ownership over outcomes, even for systems you don't own
  • Excellent communication and collaboration skills—you'll be partnering across the entire company
  • Diverse experience—the team's strength comes from people who've built product stacks, scaled databases, run massive distributed systems, and everything in between

Strong candidates may also have:

  • Been an SRE, Production Engineer, or in similar reliability-focused roles on large-scale systems
  • Experience operating large-scale model serving or training infrastructure (>1000 GPUs)
  • Experience with one or more ML hardware accelerators (GPUs, TPUs, Trainium)
  • Understanding of ML-specific networking optimizations like RDMA and InfiniBand
  • Expertise in AI-specific observability tools and frameworks
  • Experience with chaos engineering and systematic resilience testing
  • Contributed to open-source infrastructure or ML tooling

Education and Experience: Bachelor's degree or equivalent combination of education, training, and experience in a field relevant to the role. Staff-level experience is expected.

Location: London, with a hybrid policy requiring at least 25% time in office. Some roles may require more office presence.

Visa Sponsorship: The client does sponsor visas and will make every reasonable effort to secure sponsorship for successful candidates, working with immigration counsel.

How We Work With You

As Get A Job.ai's talent team, we partner exclusively with this organization to identify exceptional reliability engineering talent. When you apply through our platform at getajob.ai, one of our recruiters will conduct an initial screening to understand your background and match it against our client's needs. If there's a strong fit, we'll submit your profile to the client for consideration and guide you through their interview process.

Please do not attempt to contact the client directly, as all applications must be coordinated through Get A Job.ai to be considered for this confidential search.

Pay

Pay via Get A Job.ai: £325,000–£390,000 GBP annually

Equal Opportunity

Get A Job.ai is committed to fostering an inclusive recruitment process. We encourage applications from candidates of all backgrounds and experiences, even if you don't meet every qualification listed. Research shows that underrepresented groups are more likely to experience imposter syndrome—we urge you to apply if this work interests you.

Apply with Get A Job.ai

A recruiter will review your profile and submit you to the client. Do not contact the client directly.

More options

Apply with Get A Job.ai

Apply through Get A Job.ai. A recruiter will review your profile and submit you. Do not contact the client directly.

Apply through Get A Job.ai. A recruiter will review your profile and submit you.

Local insights for this role are preparing — this section updates automatically in a few seconds (or refresh).

Listing facts

  • Role Staff Software Engineer, AI Reliability Engineering
  • Employer Get A Job.ai
  • Location London
  • Type Full Time
  • Pay (from listing) Pay not listed
  • Posted September 16, 2026
  • Apply by October 16, 2026
  • Overview Full job description on this page (538 words)

Facts above come from this job record on Get A Job.AI — not copied from third-party review sites.

Limited public data for this employer

We only show facts we can ground in public sources (Wikidata, O*NET, news/discussion links, or this listing). We do not invent Glassdoor-style ratings, salaries, or testimonials when data is thin. Use the listing facts, occupation context, and related openings below while we continue researching.

Explore related openings

Keep exploring on Get A Job.ai

Not quite the right fit? Your next opportunity is a click away.

Hiring instead? Post a job and reach candidates searching right now.

Staff Software Engineer, AI Reliability Engineer… Get A Job.ai · London