Loading...

Senior Site Reliability Engineer — Token Factory (Inference Platform)

  • Company: Get A Job.ai
  • Location: London
  • Salary: Pay not listed
  • Full Time
  • London

Website Get A Job.ai

Represented by Get A Job.ai

About This Opportunity

We are representing a confidential AI cloud infrastructure organization that is seeking a Senior Site Reliability Engineer to join their inference platform team in London. This role sits at the intersection of massive-scale GPU orchestration and production AI model serving, where you'll own the reliability, performance, and observability of systems running foundation models at extreme scale.

As part of our talent team at Get A Job.ai, we're looking for an engineer who thrives on making distributed systems behave flawlessly under load and recover gracefully when things go wrong. You'll be working on infrastructure that serves text, vision, audio, and multimodal AI workloads to users around the world.

Responsibilities

  • Design and refine telemetry pipelines—metrics, logs, and traces—transforming hundreds of terabytes of signal into actionable insights
  • Tune Kubernetes autoscalers to maximize GPU efficiency and resource utilization
  • Craft Terraform modules that embed resilience into every cluster deployment
  • Harden request-routing and retry logic to ensure transient failures remain invisible to end users
  • Build automation and runbooks that detect, isolate, and remediate incidents in minutes
  • Drive post-mortem culture and implement preventive measures to avoid recurrence
  • Collaborate with software engineers to integrate reliability as a core feature of the platform
  • Scale the platform smoothly while meeting aggressive cost and reliability targets

What We're Looking For

  • Deep expertise with Kubernetes, Prometheus, Grafana, and Terraform
  • Strong infrastructure-as-code practices and experience managing production systems
  • Proficiency scripting in Python or Bash for automation and tooling
  • Understanding of alert design, SLOs, and observability for high-throughput APIs
  • Production experience with distributed systems and knowledge of real-world failure modes
  • Experience with GPU-heavy workloads and frameworks such as vLLM, Triton, or Ray is highly desirable
  • Background in MLOps or model-hosting platforms is a strong plus
  • Passion for building self-healing systems and debugging performance from kernel to application layer
  • Authorization to work in the United Kingdom

How We Work With You

When you apply through Get A Job.ai, one of our recruiters will review your profile and conduct an initial screening conversation. If there's a strong match, we'll submit your application to our client on your behalf and guide you through their interview process. Please do not contact the client directly—all communication should flow through our team to ensure the best experience for everyone involved.

Pay

Compensation details will be discussed during the screening process with our recruiting team.

Equal Opportunity: Get A Job.ai is committed to inclusive hiring practices. We welcome applications from candidates of all backgrounds and work to ensure fair treatment throughout our recruitment process.

Apply with Get A Job.ai

A recruiter will review your profile and submit you to the client. Do not contact the client directly.

More options

Apply with Get A Job.ai

Apply through Get A Job.ai. A recruiter will review your profile and submit you. Do not contact the client directly.

Apply through Get A Job.ai. A recruiter will review your profile and submit you.

Local insights for this role are preparing — this section updates automatically in a few seconds (or refresh).

Listing facts

  • Role Senior Site Reliability Engineer — Token Factory (Inference Platform)
  • Employer Get A Job.ai
  • Location London
  • Type Full Time
  • Pay (from listing) Pay not listed
  • Posted September 15, 2026
  • Apply by October 15, 2026
  • Overview Full job description on this page (422 words)

Facts above come from this job record on Get A Job.AI — not copied from third-party review sites.

Limited public data for this employer

We only show facts we can ground in public sources (Wikidata, O*NET, news/discussion links, or this listing). We do not invent Glassdoor-style ratings, salaries, or testimonials when data is thin. Use the listing facts, occupation context, and related openings below while we continue researching.

Explore related openings

Keep exploring on Get A Job.ai

Not quite the right fit? Your next opportunity is a click away.

Hiring instead? Post a job and reach candidates searching right now.

Senior Site Reliability Engineer — Token Factory… Get A Job.ai · London