Full Time
Anywhere
Posted 1 week ago

About JetBrains

jetbrains.com

Founded 2002
Employees 51

Source: Wikipedia

JetBrains

At JetBrains, code is our passion. Ever since we started back in 2000, we have been striving to make the strongest, most effective developer tools on earth. By automating routine checks and corrections, our tools speed up production, freeing developers to grow, discover, and create.

We’re looking for a Research Engineer who will own the training stack and model architecture for our Mellum LLM family. Your job is easier said than done: make training faster, cheaper, and more stable at a large scale. You’ll profile, design, and implement changes to the training pipeline – from architecture to custom GPU kernels, as needed.

As part of our team, you will:

Be responsible for improving end-to-end performance for multi-node LLM pre-training and post-training pipelines.
Profile hotspots (Nsight Systems/Compute, NVTX) and fix them using compute/comm overlap, kernel fusion, scheduling, etc.
Design and evaluate architecture choices (depth/width, attention variants including GQA/MQA/MLA/Flash-style, RoPE scaling/NTK, and MoE routing and load-balancing).
Implement custom ops (Triton and/or CUDA C++), integrate via PyTorch extensions, and upstream when possible.
Push memory/perf levers: FSDP/ZeRO, activation checkpointing, FP8/TE, tensor/pipeline/sequence/expert parallelism, NCCL tuning.
Harden large runs by building elastic and fault-tolerant training setups, ensuring robust checkpointing, strengthening reproducibility, and improving resilience to preemption.
Keep the data path fast using streaming and sharded data loaders and tokenizer pipelines, as well as improve overall throughput and cache efficiency.
Define the right metrics, build dashboards, and deliver steady improvements.
Run both pre-training and post-training (including SFT, RLHF, and GRPO-style methods) efficiently across sizable clusters.

We’ll be happy to bring you on board if you have:

Strong PyTorch and PyTorch Distributed experience, having run multi-node jobs with tens to hundreds of GPUs.
Hands-on experience with Megatron-LM/Megatron-Core/NeMo, DeepSpeed, or serious FSDP/ZeRO expertise.
Real profiling expertise (Nsight Systems/Compute, nvprof) and experience with NVTX-instrumented workflows.
GPU programming skills with Triton and/or CUDA, and the ability to write, test, and debug kernels.
A solid understanding of NCCL collectives, as well as topology and fabric effects (IB/RoCE), and how they show up in traces.

Our ideal candidate would have experience with:

FlashAttention-2 and 3, CUTLASS and CuTe, TransformerEngine and FP8, Inductor, AOTAutograd, and torch.compile.
MoE at scale (expert parallel, router losses, capacity management) and long-context tricks (ALiBi/YaRN/NTK scaling).
Kubernetes or SLURM at scale, placement and affinity tuning, as well as AWS, GCP, and Azure GPU fleets.
Web-scale data plumbing (streaming datasets, Parquet and TFRecord, tokenizer perf), eval harnesses, and benchmarking.
Safety and post-training methods, such as DPO, ORPO, GRPO, and reward models.
Inference ecosystems such as vLLM and paged KV.

#LI-KP1

We are an equal opportunity employer

We know great ideas can come from anyone, anywhere. That’s why we do our best to create an open and inclusive workplace – one that welcomes everyone regardless of their background, identity, religion, age, accessibility needs, or orientation.

We process the data provided in your job application in accordance with the Recruitment Privacy Policy.

To apply for this job please visit job-boards.eu.greenhouse.io.

About this role & career path

Working in Amsterdam

Amsterdam is the capital and largest city of the Netherlands. It has a population of 933,680 in June 2024 within the city proper, 1,457,018 in the urban area and 2,480,394 in the metropolitan area. Located in the Dutch province of North Holland, Amsterdam is colloquially referred to as the "Venice of the North", for its large number of canals, now a UNESCO World Heritage Site.

What people say about JetBrains

Why Write Python in Visual Studio? Hacker News · 2015-08-03
Why Write Python in Visual Studio? Hacker News · 2015-08-03
Programming fonts Hacker News · 2015-08-03
Ask HN: Has anyone here programmed in Kotlin? What do you think about it? Hacker News · 2015-07-28

More jobs at JetBrains

Keep exploring on Get A Job.ai

Not quite the right fit? Your next opportunity is a click away.

Browse all jobs
More jobs by category
Remote jobs you can do from anywhere
Research typical pay for this role
Set a job alert so new matches reach you first
Upload your resume to apply faster

Hiring instead? Post a job and reach candidates searching right now.

Get A Job.ai

Research Engineer (LLM Training and Performance)

About JetBrains

As part of our team, you will:

We’ll be happy to bring you on board if you have:

Our ideal candidate would have experience with:

About this role & career path

Working in Amsterdam

What people say about JetBrains

Recent news

More jobs at JetBrains

Keep exploring on Get A Job.ai

Nursing Assistant

Key Account Manager / Technischer Vertrieb / B2B SalesDefence, Elektronik, Optronik (m/w/d)

Business Development & Sales Manager:in (m/w/d)

Backend Developer Fintech (m/w/d)

Executive Assistant & Senior Project Manager (m/w/d)

Praktikum Financial Consulting (m/w/d) – Berlin

Werkstudent Finance & Unternehmertum (m/w/d) – Berlin

Technischer Projektmanager Softwareentwicklung (m/w/d)

Top-Karrierechance im Private Banking

Workday Financials Analyst Consultant (Global Delivery Center)