- Company: Get A Job.ai
- Location: London
- Salary: Pay not listed
Website Get A Job.ai
Represented by Get A Job.ai
Responsibilities
We are representing a confidential AI technology organization in their search for a Staff Software Engineer to join their inference infrastructure team in London. In this role, you will:
- Design, build, and maintain distributed systems that serve large language models to millions of users worldwide
- Develop resilient, flexible systems that adapt in real time to real-world events
- Develop intelligent request routing, load balancing, and traffic management systems across thousands of accelerators
- Maximize compute efficiency across the fleet by autoscaling and orchestrating production, research, and experimental workloads
- Build and operate production-grade deployment pipelines for releasing new models to users
- Provide high-performance inference infrastructure that enables researchers to develop next-generation models
- Integrate new AI accelerator platforms and support inference for new model architectures
Representative projects include designing intelligent routing algorithms that optimize request distribution across many accelerators in different environments, autoscaling compute fleets to dynamically match supply with demand, building production-grade deployment pipelines, contributing to new inference features, supporting inference for new model architectures, analyzing observability data to tune performance based on real-world production workloads, and managing multi-region deployments and geographic routing for global customers.
What We're Looking For
Required qualifications:
- Proficiency in Python or Rust
- Software engineering experience building and operating distributed systems in production
- Working knowledge of containerized infrastructure (e.g., Kubernetes) and at least one major cloud platform (AWS, GCP, or Azure)
- Results-oriented, with a bias towards flexibility and impact
- Willingness to pick up slack, even if it goes outside your job description
- Desire to learn more about machine learning systems and infrastructure
- Thrive in environments where technical excellence directly drives both business results and research breakthroughs
- Care about the societal impacts of your work
- Bachelor's degree or an equivalent combination of education, training, and/or experience in a field relevant to the role
Preferred qualifications:
- Significant experience with high-performance, large-scale distributed systems
- Experience implementing and deploying machine learning systems at scale
- Experience building load balancing, request routing, or traffic management systems
- Familiarity with LLM inference optimization, batching, and caching strategies
- Deep experience operating Kubernetes and cloud infrastructure at scale
- Experience with AI accelerator platforms (GPUs, TPUs, or emerging hardware)
The role is based in London with a hybrid policy requiring at least 25% time in office. Visa sponsorship may be available for qualified candidates.
How We Work With You
Our talent team at Get A Job.ai partners with leading technology organizations to connect exceptional engineers with transformative opportunities. When you apply through our platform:
- A dedicated recruiter from Get A Job.ai will review your application and conduct an initial screening
- If there's a strong match, we'll submit your profile to our client for consideration
- We'll guide you through each stage of the interview process and provide feedback along the way
- Please do not contact the client directly—all communication should flow through Get A Job.ai to ensure the best candidate experience
Pay
This is a full-time, permanent position. Compensation details will be discussed during the screening process with our recruiting team.
Equal Employment Opportunity: Get A Job.ai is committed to creating a diverse and inclusive recruitment process. We encourage applications from candidates of all backgrounds and experiences.
Apply with Get A Job.ai
A recruiter will review your profile and submit you to the client. Do not contact the client directly.
Apply through Get A Job.ai. A recruiter will review your profile and submit you. Do not contact the client directly.
Apply through Get A Job.ai. A recruiter will review your profile and submit you.
Local insights for this role are preparing — this section updates automatically in a few seconds (or refresh).
Listing facts
- Role Staff Software Engineer, Inference
- Employer Get A Job.ai
- Location London
- Type Full Time
- Pay (from listing) Pay not listed
- Posted September 16, 2026
- Apply by October 16, 2026
- Overview Full job description on this page (523 words)
Facts above come from this job record on Get A Job.AI — not copied from third-party review sites.
Limited public data for this employer
We only show facts we can ground in public sources (Wikidata, O*NET, news/discussion links, or this listing). We do not invent Glassdoor-style ratings, salaries, or testimonials when data is thin. Use the listing facts, occupation context, and related openings below while we continue researching.
Explore related openings
Keep exploring on Get A Job.ai
Not quite the right fit? Your next opportunity is a click away.
- Browse all jobs
- More jobs by category
- Remote jobs you can do from anywhere
- Research typical pay for this role
- Set a job alert so new matches reach you first
- Upload your resume to apply faster
Hiring instead? Post a job and reach candidates searching right now.
