- Company: Get A Job.ai
- Location: Berlin
- Salary: Pay not listed
Website Get A Job.ai
Represented by Get A Job.ai
About This Opportunity
We are representing a confidential AI cloud infrastructure organization seeking a Senior ML Engineer to join their high-performance inference and fine-tuning platform team in Berlin. This role focuses on maximizing GPU performance for foundation model deployments at massive scale.
Our client operates one of the world's largest GPU clouds, running tens of thousands of GPUs. You'll work on pushing foundation models to their hardware limits, optimizing throughput, minimizing latency, and reducing cost-per-token across production workloads.
Responsibilities
- Identify and resolve LLM inference bottlenecks to drive production performance improvements across diverse model architectures at scale
- Implement novel speculative decoding architectures and optimize components for various LLM designs including dense, MoE, autoregressive, and parallel architectures
- Contribute to open-source inference engines and optimization frameworks
- Design and productionize low-precision training and inference pipelines (FP8, NVFP4/MXFP4) with measurable gains in throughput and cost-efficiency
- Profile and optimize GPU workloads using industry-standard tools to maximize hardware utilization
- Collaborate with engineering teams to deliver high-impact optimizations in a fast-paced environment
What We're Looking For
Required qualifications:
- Profound understanding of machine learning theoretical foundations and transformer architecture
- Hands-on experience profiling GPU workloads using Nsight, PyTorch profiler, or similar tools
- Deep understanding of GPU memory hierarchy and compute/memory tradeoffs
- Strong familiarity with modern LLM concepts: MHA, RoPE, KV-cache, Flash Attention, and quantization techniques
- Understanding of performance aspects in large neural network training including sharding strategies, custom kernels, and hardware features
- Strong software engineering skills with Python as primary language
- Deep experience with modern deep learning frameworks (PyTorch, TensorFlow, JAX)
- Proficiency in contemporary software engineering practices including CI/CD, version control, and unit testing
- Strong communication and leadership abilities
Nice to have:
- Experience with open-source inference engines such as vLLM, SGLang, or TensorRT-LLM, including code contributions
- Experience with kernel languages or DSLs: Triton, Cute, CUTLASS, or CUDA
- Track record of building and shipping products in dynamic startup environments
- Experience developing large distributed systems or high-load web services
- Open-source contributions that demonstrate engineering excellence
- Excellent English communication and technical writing skills
How We Work With You
When you apply through Get A Job.ai, our talent team will review your background and schedule an initial screening call to discuss your experience with GPU optimization, LLM inference, and relevant technical projects. If there's a strong match, we'll submit your profile to our client for consideration. Please apply exclusively through our platform—direct contact with the client is not part of this process.
Our client offers competitive compensation, significant learning opportunities working with cutting-edge AI infrastructure, flexibility and ownership in your work, and the chance to make meaningful impact on large-scale GPU deployments serving the global AI community.
Pay
Compensation details will be discussed during the screening process based on your experience level and qualifications.
Equal Opportunity: Get A Job.ai is committed to inclusive hiring practices. We welcome applications from candidates of all backgrounds and provide equal employment opportunities regardless of race, color, religion, sex, national origin, age, disability, or any other protected characteristic.
Apply with Get A Job.ai
A recruiter will review your profile and submit you to the client. Do not contact the client directly.
Apply through Get A Job.ai. A recruiter will review your profile and submit you. Do not contact the client directly.
Apply through Get A Job.ai. A recruiter will review your profile and submit you.
Local insights for this role are preparing — this section updates automatically in a few seconds (or refresh).
Listing facts
- Role Senior ML Engineer (Token Factory)
- Employer Get A Job.ai
- Location Berlin
- Type Full Time
- Pay (from listing) Pay not listed
- Posted September 15, 2026
- Apply by October 16, 2026
- Country Germany
- Overview Full job description on this page (495 words)
Facts above come from this job record on Get A Job.AI — not copied from third-party review sites.
Limited public data for this employer
We only show facts we can ground in public sources (Wikidata, O*NET, news/discussion links, or this listing). We do not invent Glassdoor-style ratings, salaries, or testimonials when data is thin. Use the listing facts, occupation context, and related openings below while we continue researching.
Explore related openings
Occupation family: ML Engineer
Keep exploring on Get A Job.ai
Not quite the right fit? Your next opportunity is a click away.
- Browse all jobs
- More jobs by category
- Remote jobs you can do from anywhere
- Research typical pay for this role
- Set a job alert so new matches reach you first
- Upload your resume to apply faster
Hiring instead? Post a job and reach candidates searching right now.
