Loading...

ML Engineer, Infrastructure

  • Company: Get A Job.ai
  • Location: Berlin
  • Salary: Pay not listed
  • Full Time
  • Berlin

Website Get A Job.ai

Represented by Get A Job.ai

About This Opportunity

Our talent team at Get A Job.ai is representing a confidential AI research organization building foundation models for structured data. This is a high-impact infrastructure role where your decisions directly affect millions in compute spend and research velocity.

The team operates multi-cluster GPU infrastructure at scale, spending tens of millions annually on training compute. They're scaling from single-cluster operations to multi-provider architecture and evaluating next-generation hardware as it becomes available. You'll own the full infrastructure stack and work directly with researchers to make architectural decisions that accelerate their work.

Responsibilities

  • Own and evolve multi-cluster GPU infrastructure running Slurm on cloud platforms, with expansion to multi-provider environments and new hardware generations
  • Drive GPU utilization and training throughput through profiling, memory optimization, and systems-level debugging of distributed training workloads
  • Architect next-generation infrastructure including multi-cluster orchestration, provider diversification, and capacity planning against growing compute demands
  • Build developer productivity tooling: CI pipelines, experiment tracking, model registry, data processing systems, and internal tools that maintain research iteration speed
  • Manage compute budget with deep understanding of cost per FLOP across providers and hardware configurations
  • Work with tech stack including Slurm, GCP, Docker, wandb, GitHub Actions, uv, PyTorch, and Triton

What We're Looking For

  • 3+ years building and operating production GPU infrastructure or distributed training systems at scale in AI labs, ML startups, or HPC environments
  • Deep hands-on experience with Slurm and cluster management, including debugging scheduling failures and optimizing utilization across multi-tenant GPU workloads
  • Expert-level systems thinking around memory bandwidth, GPU profiling, and hardware-level reasoning beyond configuration
  • Strong Python skills and genuine fluency with PyTorch internals sufficient to profile training runs and identify bottlenecks
  • Track record of infrastructure decisions that measurably improved training throughput or cost efficiency
  • Strong AI tooling skills with fluency in tools like Claude Code or Cursor

Bonus qualifications:

  • Experience operating at tens-of-millions-scale GPU spend
  • Multi-cloud or hybrid HPC/cloud infrastructure experience
  • Triton, CUDA, or custom kernel experience
  • Experience scaling from single-cluster to multi-cluster orchestration
  • Background building experiment tracking, model registry, or ML pipeline tooling

How We Work With You

When you apply through Get A Job.ai, our recruiting team will conduct an initial screening to understand your background and ensure strong alignment with the role. We'll then submit qualified candidates to our client for their review. Please do not attempt to contact the client directly, as all communication should flow through our team to ensure a smooth process.

The position is based in Berlin with teams also in other European and US locations. The organization values in-person collaboration for complex technical work, though exceptional remote candidates may be considered with frequent travel expectations.

Pay

Compensation details will be discussed during the screening process based on experience and qualifications.

Equal Opportunity: Get A Job.ai is committed to inclusive hiring practices. We welcome applications from candidates of all backgrounds and provide equal opportunities regardless of gender, sexual orientation, origin, disability, or any other protected characteristic.

Apply with Get A Job.ai

A recruiter will review your profile and submit you to the client. Do not contact the client directly.

More options

Apply with Get A Job.ai

Apply through Get A Job.ai. A recruiter will review your profile and submit you. Do not contact the client directly.

Apply through Get A Job.ai. A recruiter will review your profile and submit you.

Local insights for this role are preparing — this section updates automatically in a few seconds (or refresh).

Listing facts

  • Role ML Engineer, Infrastructure
  • Employer Get A Job.ai
  • Location Berlin
  • Type Full Time
  • Pay (from listing) Pay not listed
  • Posted September 21, 2026
  • Apply by October 22, 2026
  • Country Germany
  • Overview Full job description on this page (486 words)

Facts above come from this job record on Get A Job.AI — not copied from third-party review sites.

Limited public data for this employer

We only show facts we can ground in public sources (Wikidata, O*NET, news/discussion links, or this listing). We do not invent Glassdoor-style ratings, salaries, or testimonials when data is thin. Use the listing facts, occupation context, and related openings below while we continue researching.

Explore related openings

Keep exploring on Get A Job.ai

Not quite the right fit? Your next opportunity is a click away.

Hiring instead? Post a job and reach candidates searching right now.

ML Engineer, Infrastructure Get A Job.ai · Berlin