- Company: Get A Job.ai
- Location: Paris
- Salary: Pay not listed
Website Get A Job.ai
Represented by Get A Job.ai
Responsibilities
We are representing a confidential AI safety organization seeking a Multimodal ML Engineer to join their Paris-based team. In this role, you will:
- Train and fine-tune large-scale multimodal models including vision-language, audio, and speech systems from scratch and from pretrained checkpoints
- Extend models across modalities: image understanding, video temporal modeling, long-context processing, and streaming audio
- Design and execute experiments around architecture changes, data composition, and training recipes
- Build and maintain multimodal data pipelines from raw images, video, and audio recordings to training-ready datasets, including synthetic data generation
- Train and optimize Mixture of Experts (MoE) architectures for efficient multimodal inference
- Develop alignment pipelines including supervised fine-tuning, direct preference optimization, and reward modeling across multiple modalities
- Optimize models for production through quantization, distillation, batching, and low-latency serving
- Deploy models end-to-end from research checkpoint to production systems
- Define evaluation metrics and benchmarks for visual QA, spatial reasoning, video comprehension, and speech understanding
What We're Looking For
Our client needs someone with:
- 3+ years of experience training large-scale deep learning models in multimodal domains (vision-language, audio, speech, or acoustic)
- Strong PyTorch skills with hands-on distributed training experience using DeepSpeed, FSDP, or similar frameworks
- Deep experience with multimodal architectures and understanding of how vision/audio encoders, projectors, and language models integrate
- Hands-on experience with RLHF and alignment techniques for multimodal systems, including GRPO, DPO, and reward modeling beyond text-only applications
- Experience with video and/or audio sequence modeling: temporal modeling, long-context processing, efficient attention mechanisms, and streaming inference
- Proven track record of shipping models to production, meeting latency targets, and optimizing inference performance
- Comfort with large-scale multimodal dataset curation including image-text pairs, video-instruction data, audio preprocessing, augmentation, and synthetic data generation
- Familiarity with MoE architectures and their tradeoffs for multimodal workloads
- Strong engineering fundamentals: clean code, version control, testing, and documentation
Bonus qualifications: Understanding of audio signal processing fundamentals including spectrograms, mel features, and noise reduction.
How We Work With You
Our talent team at Get A Job.ai will guide you through this confidential search:
- Apply directly through the Get A Job.ai platform with your resume and relevant portfolio or GitHub links
- A recruiter from our team will conduct an initial screening to understand your background and motivations
- We will submit qualified candidates to our client for consideration
- Please do not attempt to contact the client directly, as this is a confidential search
- Our client's interview process includes a take-home technical task, a technical interview with their Head of Applied Research, and a final conversation with leadership
Location and Benefits
This position is based in Paris with hybrid work arrangements. A relocation package is available for qualified candidates. The client offers:
- Competitive paid time off in line with local regulations
- Comprehensive medical insurance for France-based team members
- Meaningful equity package
- All necessary hardware, tools, and services
- Covered subscriptions for AI development tools
- Team off-site events twice annually
Pay
Compensation details will be discussed during the screening process with Get A Job.ai and are competitive for senior multimodal ML engineering roles in the Paris market.
Equal Opportunity: Get A Job.ai is committed to inclusive recruiting practices. We partner with clients who value diversity and welcome applicants from all backgrounds.
Apply with Get A Job.ai
A recruiter will review your profile and submit you to the client. Do not contact the client directly.
Apply through Get A Job.ai. A recruiter will review your profile and submit you. Do not contact the client directly.
Apply through Get A Job.ai. A recruiter will review your profile and submit you.
Local insights for this role are preparing — this section updates automatically in a few seconds (or refresh).
Listing facts
- Role Multimodal ML Engineer
- Employer Get A Job.ai
- Location Paris
- Type Full Time
- Pay (from listing) Pay not listed
- Posted September 6, 2026
- Apply by October 6, 2026
- Overview Full job description on this page (527 words)
Facts above come from this job record on Get A Job.AI — not copied from third-party review sites.
Limited public data for this employer
We only show facts we can ground in public sources (Wikidata, O*NET, news/discussion links, or this listing). We do not invent Glassdoor-style ratings, salaries, or testimonials when data is thin. Use the listing facts, occupation context, and related openings below while we continue researching.
Explore related openings
Keep exploring on Get A Job.ai
Not quite the right fit? Your next opportunity is a click away.
- Browse all jobs
- More jobs by category
- Remote jobs you can do from anywhere
- Research typical pay for this role
- Set a job alert so new matches reach you first
- Upload your resume to apply faster
Hiring instead? Post a job and reach candidates searching right now.
