- Company: Get A Job.ai
- Location: London
- Salary: Pay not listed
Website Get A Job.ai
Represented by Get A Job.ai
Responsibilities
We are representing a confidential cloud infrastructure organization seeking a Senior Data & MLOps Engineer in London. In this role, you will design and scale the infrastructure supporting an advanced reliability platform that processes telemetry data, derives meaningful metrics, identifies anomalies, predicts potential issues, and enables automated root cause analysis across distributed systems.
Your responsibilities will include:
- Design and implement scalable data ingestion pipelines for high-throughput streaming and telemetry data
- Build feature processing and baseline computation systems for time-series data
- Productionize machine learning models for prediction and anomaly detection
- Develop and operate low-latency services and robust offline workflows
- Architect horizontally scalable microservices with clear separation between components, leveraging orchestration technologies
- Implement monitoring and feedback loops for continuous model and signal improvement
- Collaborate with Platform, Infrastructure, and Fleet teams to integrate operational signals into monitoring and diagnostics
- Scale systems from proof-of-concept to production-grade, fleet-level deployments
- Implement scalable solutions for mitigation and structured analysis
What We're Looking For
Required qualifications:
- 7+ years of experience in data engineering, distributed systems, MLOps, or infrastructure ML roles in production environments
- Proven experience building high-throughput streaming or telemetry pipelines (Kafka, Pulsar, Kinesis, or equivalent)
- Strong experience designing time-series feature pipelines and operating large-scale observability systems
- Experience building and maintaining feature stores and ensuring offline/online feature parity
- Hands-on experience deploying ML models to production, including versioning, monitoring, rollback, and drift detection
- Experience designing scalable microservices deployed in Kubernetes-based environments
- Strong proficiency in Python and at least one systems language (Go, Rust, or C++)
- Experience working with distributed compute or training systems (NCCL, PyTorch Distributed, Spark, Ray, Slurm)
- Familiarity with GPU telemetry systems such as NVML or DCGM and hardware-level monitoring concepts
- Demonstrated experience scaling systems from proof-of-concept to production-grade deployments
Preferred qualifications:
- Experience working on GPU fleet management, hyperscale infrastructure, or AI training clusters
- Experience building anomaly detection or failure prediction systems for hardware or distributed systems
- Experience implementing distributed straggler detection or collective-level performance analysis systems
- Experience developing agentic or LLM-powered reasoning systems for diagnostics or operational intelligence
- Background in reliability engineering or SRE practices
You'll thrive in this role if:
- You love building systems that turn raw infrastructure telemetry into actionable intelligence
- You're curious about distributed systems failure modes, GPU performance pathologies, and reliability engineering at scale
- You're excited by the idea of moving from anomaly detection to prediction to autonomous root cause reasoning
- You enjoy designing platforms that protect uptime, revenue, and customer trust through proactive systems thinking
How We Work With You
Our talent team at Get A Job.ai will guide you through the entire process. When you apply through our platform, a dedicated recruiter will screen your application and conduct an initial conversation to understand your background and career goals. Once we determine there's a strong fit, we'll submit your profile to our client for consideration and coordinate all subsequent interview stages.
Please apply exclusively through Get A Job.ai. Do not attempt to contact the client organization directly, as this may disqualify your application.
Pay
Compensation details will be discussed during the screening process with our recruiting team.
Equal Opportunity
Get A Job.ai is committed to fostering an inclusive recruitment process. We welcome applications from all qualified candidates regardless of race, color, religion, sex, disability, age, sexual orientation, gender identity, national origin, veteran status, or any other protected characteristic.
Apply with Get A Job.ai
A recruiter will review your profile and submit you to the client. Do not contact the client directly.
Apply through Get A Job.ai. A recruiter will review your profile and submit you. Do not contact the client directly.
Apply through Get A Job.ai. A recruiter will review your profile and submit you.
Explore Get A Job.ai online
Working in London, UK
Weather right now in London, UK: checking… · Local time: · Air quality: · Daylight: · UV index: · Wind: · Pollen:
London is the capital and largest city of England and the United Kingdom, with a population of 9.1 million people in 2024. Its wider metropolitan area is the largest in Western Europe, with a population of 15.4 million. London stands on the River Thames in southeast England, at the head of a 50-mile (80 km) tidal estuary down to the North Sea, and has been a major settlement for nearly 2,000 years. Its ancient core and financial centre, the City of London, was founded by the Romans as Londinium and has retained its medieval boundaries. The City of Westminster, to the west of the City of London
England is a country that is part of the United Kingdom. It is located on the island of Great Britain, of which it covers about 62%, and more than 100 smaller adjacent islands. England shares a land border with Scotland to the north and another land border with Wales to the west, and is surrounded by the North Sea to the east, the English Channel to the south, the Celtic Sea to the south-west, and
Nearby green space: 12 parks within 1.5km — closest is Whitehall Garden (364m). via OpenStreetMap
Nearest public transit: Charing Cross (station, 41m). via OpenStreetMap
- Elevation 18m (59 ft)
Source: Wikipedia (state)
Job details above are provided by the employer/source. The sections on this page are compiled from public data sources with AI assistance.
Accommodations: if you need a workplace accommodation to apply for or perform this job, see ADA.gov or EEOC.gov for guidance on your rights and how to request one.
Listing facts
- Role Senior Data & MLOps Engineer
- Employer Get A Job.ai
- Location London
- Type Full Time
- Pay (from listing) Pay not listed
- Posted September 23, 2026
- Apply by October 24, 2026
- Overview Full job description on this page (551 words)
Facts above come from this job record on Get A Job.AI — not copied from third-party review sites.
Limited public data for this employer
We only show facts we can ground in public sources (Wikidata, O*NET, news/discussion links, or this listing). We do not invent Glassdoor-style ratings, salaries, or testimonials when data is thin. Use the listing facts, occupation context, and related openings below while we continue researching.
Explore related openings
Keep exploring on Get A Job.ai
Not quite the right fit? Your next opportunity is a click away.
- Browse all jobs
- More jobs by category
- Remote jobs you can do from anywhere
- Research typical pay for this role
- Set a job alert so new matches reach you first
- Upload your resume to apply faster
Hiring instead? Post a job and reach candidates searching right now.
