Loading...

Multimodal ML Engineer

  • Company: Get A Job.ai
  • Location: Paris
  • Salary: Pay not listed
  • Full Time
  • Paris

Website Get A Job.ai

Represented by Get A Job.ai

Responsibilities

We are representing a confidential AI safety organization seeking a Multimodal ML Engineer to join their Paris-based team. In this role, you will:

  • Train and fine-tune large-scale multimodal models including vision-language, audio, and speech systems from scratch and from pretrained checkpoints
  • Extend models across modalities: image understanding, video temporal modeling, long-context processing, and streaming audio
  • Design and execute experiments around architecture changes, data composition, and training recipes
  • Build and maintain multimodal data pipelines from raw images, video, and audio recordings to training-ready datasets, including synthetic data generation
  • Train and optimize Mixture of Experts (MoE) architectures for efficient multimodal inference
  • Develop alignment pipelines including supervised fine-tuning, direct preference optimization, and reward modeling across multiple modalities
  • Optimize models for production through quantization, distillation, batching, and low-latency serving
  • Deploy models end-to-end from research checkpoint to production systems
  • Define evaluation metrics and benchmarks for visual QA, spatial reasoning, video comprehension, and speech understanding

What We're Looking For

Our client needs someone with:

  • 3+ years of experience training large-scale deep learning models in multimodal domains (vision-language, audio, speech, or acoustic)
  • Strong PyTorch skills with hands-on distributed training experience using DeepSpeed, FSDP, or similar frameworks
  • Deep experience with multimodal architectures and understanding of how vision/audio encoders, projectors, and language models integrate
  • Hands-on experience with RLHF and alignment techniques for multimodal systems, including GRPO, DPO, and reward modeling beyond text-only applications
  • Experience with video and/or audio sequence modeling: temporal modeling, long-context processing, efficient attention mechanisms, and streaming inference
  • Proven track record of shipping models to production, meeting latency targets, and optimizing inference performance
  • Comfort with large-scale multimodal dataset curation including image-text pairs, video-instruction data, audio preprocessing, augmentation, and synthetic data generation
  • Familiarity with MoE architectures and their tradeoffs for multimodal workloads
  • Strong engineering fundamentals: clean code, version control, testing, and documentation

Bonus qualifications: Understanding of audio signal processing fundamentals including spectrograms, mel features, and noise reduction.

How We Work With You

Our talent team at Get A Job.ai will guide you through this confidential search:

  • Apply directly through the Get A Job.ai platform with your resume and relevant portfolio or GitHub links
  • A recruiter from our team will conduct an initial screening to understand your background and motivations
  • We will submit qualified candidates to our client for consideration
  • Please do not attempt to contact the client directly, as this is a confidential search
  • Our client's interview process includes a take-home technical task, a technical interview with their Head of Applied Research, and a final conversation with leadership

Location and Benefits

This position is based in Paris with hybrid work arrangements. A relocation package is available for qualified candidates. The client offers:

  • Competitive paid time off in line with local regulations
  • Comprehensive medical insurance for France-based team members
  • Meaningful equity package
  • All necessary hardware, tools, and services
  • Covered subscriptions for AI development tools
  • Team off-site events twice annually

Pay

Compensation details will be discussed during the screening process with Get A Job.ai and are competitive for senior multimodal ML engineering roles in the Paris market.

Equal Opportunity: Get A Job.ai is committed to inclusive recruiting practices. We partner with clients who value diversity and welcome applicants from all backgrounds.

Apply with Get A Job.ai

A recruiter will review your profile and submit you to the client. Do not contact the client directly.

More options

Apply with Get A Job.ai

Apply through Get A Job.ai. A recruiter will review your profile and submit you. Do not contact the client directly.

Apply through Get A Job.ai. A recruiter will review your profile and submit you.

Local insights for this role are preparing — this section updates automatically in a few seconds (or refresh).

Listing facts

  • Role Multimodal ML Engineer
  • Employer Get A Job.ai
  • Location Paris
  • Type Full Time
  • Pay (from listing) Pay not listed
  • Posted September 6, 2026
  • Apply by October 6, 2026
  • Overview Full job description on this page (527 words)

Facts above come from this job record on Get A Job.AI — not copied from third-party review sites.

Limited public data for this employer

We only show facts we can ground in public sources (Wikidata, O*NET, news/discussion links, or this listing). We do not invent Glassdoor-style ratings, salaries, or testimonials when data is thin. Use the listing facts, occupation context, and related openings below while we continue researching.

Explore related openings

Keep exploring on Get A Job.ai

Not quite the right fit? Your next opportunity is a click away.

Hiring instead? Post a job and reach candidates searching right now.

Multimodal ML Engineer Get A Job.ai · Paris