Loading...

ML Infrastructure Engineer

  • Company: Get A Job.ai
  • Location: Paris
  • Salary: Pay not listed
  • Full Time
  • Paris

Website Get A Job.ai

Represented by Get A Job.ai

About This Opportunity

We are representing a well-funded AI safety organization in Paris that is building critical infrastructure for LLM training, reinforcement learning, and agentic systems at scale. Our client processes over 100M API calls monthly and maintains their own fine-tuned models optimized for performance and cost.

As an ML Infrastructure Engineer, you'll build the systems that power post-training pipelines, evaluation frameworks, and inference stacks. You'll work at the intersection of research and production, where infrastructure decisions directly impact model quality and learning dynamics.

Responsibilities

  • Design and build robust RL and post-training pipelines, including data control systems that govern training data flow, rollouts, filtering, and policy updates
  • Optimize training and inference end-to-end for high throughput across networking, memory, compute scheduling, data loading, storage, and checkpointing
  • Investigate how infrastructure choices affect learning dynamics, evaluation quality, model behavior, and training stability
  • Build infrastructure for model iteration: experiment tracking, artifacts management, evaluation dashboards, failure inspection, and reproducibility tools
  • Develop and improve agentic development environments including coding-agent harnesses, browser/tool integrations, runtime sandboxes, and multi-agent orchestration
  • Collaborate closely with researchers to plan infrastructure improvements, discuss tradeoffs, and maintain context across the team

What We're Looking For

Required:

  • Experience designing, building, or maintaining distributed RL/post-training systems at scale, with deep understanding of rollouts, replay buffers, reward signals, data filtering, and evaluation loops
  • Proficiency with deep learning frameworks such as PyTorch or JAX
  • Strong Python skills including concurrency, asynchronous programming, multiprocessing, and performance optimization
  • Ability to debug distributed GPU workloads across CUDA runtime, container runtime, NCCL or equivalent communication layers, networking, and storage
  • Experience with profiling tools such as py-spy, PyTorch profiler, Nsight, perf, or custom instrumentation
  • Familiarity with inference stacks like vLLM, SGLang, TensorRT-LLM, or custom serving infrastructure
  • Ability to reason from system metrics back to model behavior and understand how infrastructure affects learning
  • Strong ownership mindset with ability to take ambiguous problems from concept to production

Preferred:

  • Public builder footprint: open-source contributions to RL, distributed ML, LLM training, inference, or agent infrastructure
  • Experience in high-bar AI infrastructure or research environments
  • Custom training framework development: distributed training, fine-tuning pipelines, schedulers, checkpointing
  • Experience with agentic coding systems in production development workflows
  • GPU cluster orchestration experience with Kubernetes, Slurm, Ray, or custom schedulers
  • Deep networking experience with NCCL, UCX, RDMA, InfiniBand, or RoCE
  • Systems-level programming experience in Rust, C++, CUDA, or Go

How We Work With You

Our talent team at Get A Job.ai partners with leading AI organizations to connect exceptional infrastructure engineers with high-impact opportunities. When you apply through our platform, a recruiter will conduct an initial screening to understand your background and confirm fit. We then coordinate your submission to our client and support you throughout their interview process.

Please do not contact the employer directly. All applications and communications should go through Get A Job.ai to ensure proper representation and support.

Interview Process (conducted by our client):

  • Introductory call (25 min)
  • Take-home technical task
  • Technical interview with engineering leadership (60 min)
  • Final conversation with executive team (45 min)

Location & Benefits

Location: Paris (hybrid) with relocation package available, or London (remote work available; relocation support not currently available for London-based candidates)

Benefits include:

  • Meaningful equity package
  • Comprehensive medical insurance (France-based team members)
  • Paid time off in line with local regulations
  • All necessary hardware, tools, and services
  • Covered subscriptions for AI agents and IDEs
  • Team off-sites twice annually

Pay

Compensation details will be discussed during the screening process and are competitive for senior ML infrastructure roles in the Paris/London market.

Get A Job.ai is an equal opportunity recruiter. We welcome applications from candidates of all backgrounds and do not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status.

Apply with Get A Job.ai

A recruiter will review your profile and submit you to the client. Do not contact the client directly.

More options

Apply with Get A Job.ai

Apply through Get A Job.ai. A recruiter will review your profile and submit you. Do not contact the client directly.

Apply through Get A Job.ai. A recruiter will review your profile and submit you.

Local insights for this role are preparing — this section updates automatically in a few seconds (or refresh).

Listing facts

  • Role ML Infrastructure Engineer
  • Employer Get A Job.ai
  • Location Paris
  • Type Full Time
  • Pay (from listing) Pay not listed
  • Posted September 6, 2026
  • Apply by October 6, 2026
  • Overview Full job description on this page (618 words)

Facts above come from this job record on Get A Job.AI — not copied from third-party review sites.

Limited public data for this employer

We only show facts we can ground in public sources (Wikidata, O*NET, news/discussion links, or this listing). We do not invent Glassdoor-style ratings, salaries, or testimonials when data is thin. Use the listing facts, occupation context, and related openings below while we continue researching.

Explore related openings

Keep exploring on Get A Job.ai

Not quite the right fit? Your next opportunity is a click away.

Hiring instead? Post a job and reach candidates searching right now.

ML Infrastructure Engineer Get A Job.ai · Paris