- Company: Get A Job.ai
- Location: Paris
- Salary: Pay not listed
Website Get A Job.ai
Represented by Get A Job.ai
About This Opportunity
We are representing a well-funded AI safety organization in Paris that is building critical infrastructure for LLM training, reinforcement learning, and agentic systems at scale. Our client processes over 100M API calls monthly and maintains their own fine-tuned models optimized for performance and cost.
As an ML Infrastructure Engineer, you'll build the systems that power post-training pipelines, evaluation frameworks, and inference stacks. You'll work at the intersection of research and production, where infrastructure decisions directly impact model quality and learning dynamics.
Responsibilities
- Design and build robust RL and post-training pipelines, including data control systems that govern training data flow, rollouts, filtering, and policy updates
- Optimize training and inference end-to-end for high throughput across networking, memory, compute scheduling, data loading, storage, and checkpointing
- Investigate how infrastructure choices affect learning dynamics, evaluation quality, model behavior, and training stability
- Build infrastructure for model iteration: experiment tracking, artifacts management, evaluation dashboards, failure inspection, and reproducibility tools
- Develop and improve agentic development environments including coding-agent harnesses, browser/tool integrations, runtime sandboxes, and multi-agent orchestration
- Collaborate closely with researchers to plan infrastructure improvements, discuss tradeoffs, and maintain context across the team
What We're Looking For
Required:
- Experience designing, building, or maintaining distributed RL/post-training systems at scale, with deep understanding of rollouts, replay buffers, reward signals, data filtering, and evaluation loops
- Proficiency with deep learning frameworks such as PyTorch or JAX
- Strong Python skills including concurrency, asynchronous programming, multiprocessing, and performance optimization
- Ability to debug distributed GPU workloads across CUDA runtime, container runtime, NCCL or equivalent communication layers, networking, and storage
- Experience with profiling tools such as py-spy, PyTorch profiler, Nsight, perf, or custom instrumentation
- Familiarity with inference stacks like vLLM, SGLang, TensorRT-LLM, or custom serving infrastructure
- Ability to reason from system metrics back to model behavior and understand how infrastructure affects learning
- Strong ownership mindset with ability to take ambiguous problems from concept to production
Preferred:
- Public builder footprint: open-source contributions to RL, distributed ML, LLM training, inference, or agent infrastructure
- Experience in high-bar AI infrastructure or research environments
- Custom training framework development: distributed training, fine-tuning pipelines, schedulers, checkpointing
- Experience with agentic coding systems in production development workflows
- GPU cluster orchestration experience with Kubernetes, Slurm, Ray, or custom schedulers
- Deep networking experience with NCCL, UCX, RDMA, InfiniBand, or RoCE
- Systems-level programming experience in Rust, C++, CUDA, or Go
How We Work With You
Our talent team at Get A Job.ai partners with leading AI organizations to connect exceptional infrastructure engineers with high-impact opportunities. When you apply through our platform, a recruiter will conduct an initial screening to understand your background and confirm fit. We then coordinate your submission to our client and support you throughout their interview process.
Please do not contact the employer directly. All applications and communications should go through Get A Job.ai to ensure proper representation and support.
Interview Process (conducted by our client):
- Introductory call (25 min)
- Take-home technical task
- Technical interview with engineering leadership (60 min)
- Final conversation with executive team (45 min)
Location & Benefits
Location: Paris (hybrid) with relocation package available, or London (remote work available; relocation support not currently available for London-based candidates)
Benefits include:
- Meaningful equity package
- Comprehensive medical insurance (France-based team members)
- Paid time off in line with local regulations
- All necessary hardware, tools, and services
- Covered subscriptions for AI agents and IDEs
- Team off-sites twice annually
Pay
Compensation details will be discussed during the screening process and are competitive for senior ML infrastructure roles in the Paris/London market.
Get A Job.ai is an equal opportunity recruiter. We welcome applications from candidates of all backgrounds and do not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status.
Apply with Get A Job.ai
A recruiter will review your profile and submit you to the client. Do not contact the client directly.
Apply through Get A Job.ai. A recruiter will review your profile and submit you. Do not contact the client directly.
Apply through Get A Job.ai. A recruiter will review your profile and submit you.
Local insights for this role are preparing — this section updates automatically in a few seconds (or refresh).
Listing facts
- Role ML Infrastructure Engineer
- Employer Get A Job.ai
- Location Paris
- Type Full Time
- Pay (from listing) Pay not listed
- Posted September 6, 2026
- Apply by October 6, 2026
- Overview Full job description on this page (618 words)
Facts above come from this job record on Get A Job.AI — not copied from third-party review sites.
Limited public data for this employer
We only show facts we can ground in public sources (Wikidata, O*NET, news/discussion links, or this listing). We do not invent Glassdoor-style ratings, salaries, or testimonials when data is thin. Use the listing facts, occupation context, and related openings below while we continue researching.
Explore related openings
Keep exploring on Get A Job.ai
Not quite the right fit? Your next opportunity is a click away.
- Browse all jobs
- More jobs by category
- Remote jobs you can do from anywhere
- Research typical pay for this role
- Set a job alert so new matches reach you first
- Upload your resume to apply faster
Hiring instead? Post a job and reach candidates searching right now.
