Loading...

Senior HPC Cluster Engineer

  • Company: Get A Job.ai
  • Location: Germany, Netherlands, United Kingdom
  • Salary: Pay not listed

Website Get A Job.ai

Represented by Get A Job.ai

About This Opportunity

Our talent team at Get A Job.ai is representing a confidential cloud infrastructure organization in their search for a Senior HPC Cluster Engineer. This role offers the opportunity to work on large-scale GPU orchestration and high-performance computing platforms that power AI and machine learning workloads at scale.

The position is based in Germany, Netherlands, or the United Kingdom and involves working with cutting-edge hardware virtualization, GPU computing, and InfiniBand networking technologies.

Responsibilities

In this role, you will:

  • Tune the performance of GPU clusters and InfiniBand networks to ensure optimal operation in HPC and GPU-based environments
  • Analyze and troubleshoot root cause issues related to GPUs and InfiniBand networks, proposing and implementing corrective actions
  • Integrate new hardware into existing infrastructure, including support for new GPU hardware through software stacks like Kubernetes, QEMU, and KVM
  • Enhance automation systems for proactive monitoring, detection, and resolution of issues in GPU and InfiniBand environments
  • Configure and manage GPU devices and InfiniBand fabrics to ensure efficient and reliable operation
  • Work closely with hardware virtualization and device emulation technologies to maintain high performance and security in multi-GPU environments

What We're Looking For

Required qualifications:

  • 5+ years of professional experience in system-level software development with focus on performance optimization and low-level programming
  • 3+ years of hands-on experience with Linux systems, including administration, troubleshooting, and performance tuning
  • In-depth understanding of server architecture, including PCIe devices, NICs, Linux OS/Kernel, and high-performance computing (HPC) systems
  • Strong proficiency in one or more performance-oriented programming languages (C/C++, Go, Python)
  • Authorization to work in the country where you apply

Preferred qualifications:

  • Experience with GPU end-to-end testing in cluster environments using InfiniBand networking
  • Proven track record of analyzing and optimizing HPC workloads such as simulations, data analysis, or AI/ML workloads
  • Familiarity with RDMA, RoCE, and InfiniBand protocols for high-performance communication
  • Background in Software-Defined Networking (SDN) and HPC cluster networking
  • Understanding of QEMU/KVM virtualization and managing virtualized environments
  • Experience with deep learning frameworks such as PyTorch and TensorFlow and their integration with HPC systems
  • Familiarity with collective communication libraries like MPI and NCCL for distributed computing

Please note that coding interviews are part of the selection process.

How We Work With You

When you apply through Get A Job.ai, our recruiting team will review your profile and conduct an initial screening. If there's a strong match, we'll submit your candidacy to our client for consideration. We ask that you do not contact the client organization directly, as all communication should flow through our team to ensure the best experience for everyone involved.

Pay

Compensation details will be discussed during the screening process and are competitive for this level of expertise and the European market.

Equal Opportunity: Get A Job.ai is committed to inclusive hiring practices. We welcome applications from candidates of all backgrounds and provide equal employment opportunities regardless of race, color, religion, sex, national origin, age, disability, or any other protected characteristic.

Apply with Get A Job.ai

A recruiter will review your profile and submit you to the client. Do not contact the client directly.

More options

Apply with Get A Job.ai

Apply through Get A Job.ai. A recruiter will review your profile and submit you. Do not contact the client directly.

Apply through Get A Job.ai. A recruiter will review your profile and submit you.

Local insights for this role are preparing — this section updates automatically in a few seconds (or refresh).

Listing facts

  • Role Senior HPC Cluster Engineer
  • Employer Get A Job.ai
  • Location Germany, Netherlands, United Kingdom
  • Type Full Time
  • Pay (from listing) Pay not listed
  • Posted September 1, 2026
  • Apply by October 2, 2026
  • Country United Kingdom
  • Overview Full job description on this page (484 words)

Facts above come from this job record on Get A Job.AI — not copied from third-party review sites.

Limited public data for this employer

We only show facts we can ground in public sources (Wikidata, O*NET, news/discussion links, or this listing). We do not invent Glassdoor-style ratings, salaries, or testimonials when data is thin. Use the listing facts, occupation context, and related openings below while we continue researching.

Explore related openings

Keep exploring on Get A Job.ai

Not quite the right fit? Your next opportunity is a click away.

Hiring instead? Post a job and reach candidates searching right now.

Senior HPC Cluster Engineer Get A Job.ai · Germany, Netherlands, United Kingdom