Loading...

Senior Machine Learning Engineer, LLM Inference Optimization

  • Company: Get A Job.ai
  • Location: London
  • Salary: Pay not listed
  • Full Time
  • London

Website Get A Job.ai

Represented by Get A Job.ai

About This Opportunity

We are representing a confidential AI cloud infrastructure organization seeking a Senior Machine Learning Engineer to join their Applied AI team in London. This is a hands-on role focused on LLM inference optimization, where you will own model and endpoint optimization from artifacts through production deployment.

In this position, you will work on complex optimization projects spanning model internals, inference engines, serving architecture, and benchmarking. Your focus will be on improving latency, throughput, memory efficiency, GPU utilization, and cost per token while maintaining model quality and reliability in production environments.

Responsibilities

  • Own optimization work for specific model families, customer endpoints, or serving backends
  • Run engine comparisons and recommend practical serving configurations for specific workloads
  • Debug model quality or performance regressions during production rollouts
  • Optimize LLM and VLM endpoints for latency, throughput, memory efficiency, GPU utilization, quality, and cost per token
  • Deploy, configure, benchmark, and extend inference engines such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, or similar systems
  • Build and productionize model-compression workflows, including quantization, quantization-aware training, distillation, low-bit serving, and accuracy recovery
  • Implement or integrate speculative decoding, draft-model approaches, KV-cache optimization, prefix caching, chunked prefill, continuous batching, and disaggregated prefill/decode serving
  • Build reproducible benchmark harnesses for TTFT, TPOT, tokens per second per GPU, p95/p99 latency, GPU memory, reliability, and cost per token
  • Partner with GPU kernel engineers and platform engineers to diagnose bottlenecks across model code, kernels, runtime, scheduler, gateway, and cluster layers
  • Write clear design docs, performance reports, rollout plans, and customer-facing technical explanations

What We're Looking For

Must-Have Requirements:

  • Strong Python and PyTorch engineering skills
  • Hands-on experience deploying or optimizing LLM, VLM, or high-throughput transformer inference systems
  • Practical knowledge of at least one modern inference stack such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, Ray Serve, KServe, or equivalent internal systems
  • Strong understanding of transformer inference bottlenecks, including KV cache, attention, memory bandwidth, batching, parallelism, and long-context serving
  • Ability to reason quantitatively about latency, throughput, quality, utilization, and cost tradeoffs
  • Strong communication skills and ability to collaborate with research, kernel, infrastructure, product, and customer teams

Nice-to-Have Qualifications:

  • Experience with quantization-aware training, post-training quantization, FP8, INT8, INT4, NVFP4, MXFP4, AWQ, GPTQ, SmoothQuant, or related techniques
  • Experience with distillation, speculative decoding, EAGLE, Medusa, multi-token prediction, or other inference acceleration methods
  • Experience with agentic workloads, including tool calling, structured outputs, streaming APIs, high concurrency, and multi-step orchestration
  • CUDA or Triton familiarity, even if the role is not primarily a kernel-engineering role
  • Open-source contributions to vLLM, SGLang, TensorRT-LLM, FlashInfer, LMCache, PyTorch, Triton, Ray, KServe, or related projects

How We Work With You

When you apply through Get A Job.ai, our talent team will review your profile and conduct an initial screening. If there's a strong match, we will submit your candidacy to our client for consideration. Please note that you should not contact the client directly—all communication and coordination will be managed through our recruiting team to ensure the best experience for everyone involved.

Pay

Compensation details will be discussed during the screening process based on experience and qualifications.

Equal Opportunity: Get A Job.ai is committed to inclusive hiring practices. We welcome applications from candidates of all backgrounds and work to ensure fair consideration throughout our recruitment process.

Apply with Get A Job.ai

A recruiter will review your profile and submit you to the client. Do not contact the client directly.

More options

Apply with Get A Job.ai

Apply through Get A Job.ai. A recruiter will review your profile and submit you. Do not contact the client directly.

Apply through Get A Job.ai. A recruiter will review your profile and submit you.

Working in London, UK

Weather right now in London, UK: checking… · Local time: · Air quality: · Daylight: · UV index: · Wind: · Pollen:

London is the capital and largest city of England and the United Kingdom, with a population of 9.1 million people in 2024. Its wider metropolitan area is the largest in Western Europe, with a population of 15.4 million. London stands on the River Thames in southeast England, at the head of a 50-mile (80 km) tidal estuary down to the North Sea, and has been a major settlement for nearly 2,000 years. Its ancient core and financial centre, the City of London, was founded by the Romans as Londinium and has retained its medieval boundaries. The City of Westminster, to the west of the City of London

England is a country that is part of the United Kingdom. It is located on the island of Great Britain, of which it covers about 62%, and more than 100 smaller adjacent islands. England shares a land border with Scotland to the north and another land border with Wales to the west, and is surrounded by the North Sea to the east, the English Channel to the south, the Celtic Sea to the south-west, and

Nearby green space: 12 parks within 1.5km — closest is Whitehall Garden (364m). via OpenStreetMap

Nearest public transit: Charing Cross (station, 41m). via OpenStreetMap

  • Elevation 18m (59 ft)

Source: Wikipedia (state)

Job details above are provided by the employer/source. The sections on this page are compiled from public data sources with AI assistance.

Accommodations: if you need a workplace accommodation to apply for or perform this job, see ADA.gov or EEOC.gov for guidance on your rights and how to request one.

Add application deadline to calendar

Listing facts

  • Role Senior Machine Learning Engineer, LLM Inference Optimization
  • Employer Get A Job.ai
  • Location London
  • Type Full Time
  • Pay (from listing) Pay not listed
  • Posted September 23, 2026
  • Apply by October 23, 2026
  • Overview Full job description on this page (539 words)

Facts above come from this job record on Get A Job.AI — not copied from third-party review sites.

Limited public data for this employer

We only show facts we can ground in public sources (Wikidata, O*NET, news/discussion links, or this listing). We do not invent Glassdoor-style ratings, salaries, or testimonials when data is thin. Use the listing facts, occupation context, and related openings below while we continue researching.

Explore related openings

Keep exploring on Get A Job.ai

Not quite the right fit? Your next opportunity is a click away.

Hiring instead? Post a job and reach candidates searching right now.

Senior Machine Learning Engineer, LLM Inference… Get A Job.ai · London