- Company: Get A Job.ai
- Location: United Kingdom
- Salary: Pay not listed
Website Get A Job.ai
Represented by Get A Job.ai
About This Opportunity
We are representing a confidential cloud infrastructure organization that is experiencing significant growth in enterprise GPU deployments. Our talent team is seeking an HPC Cluster Architect to own the full architecture lifecycle for large-scale dedicated GPU clusters—from initial customer requirements through to production deployment.
This is a senior, hands-on role based in the United Kingdom (remote) where you'll have direct ownership over cluster architecture spanning compute, networking, storage, and physical design. You'll be the technical authority translating customer needs into production-ready, commercially optimized GPU deployments.
Responsibilities
In this role, success looks like:
- Owning end-to-end cluster architecture for large-scale NVIDIA GPU deployments—from customer requirements through rack layouts, bill of materials, power and cooling design, to production handover
- Designing high-performance network fabrics across compute (InfiniBand, RDMA, NVLink/NVSwitch), storage, and WAN—defining topology, oversubscription models, and scaling strategies
- Engaging directly with OEMs and vendors to validate hardware configurations, review quotes, and ensure designs are both technically sound and commercially optimized
- Providing technical oversight during deployment and bring-up phases—supporting hardware validation, performance testing, and serving as escalation point for complex integration issues
- Acting as senior technical leader across solutions architecture, cloud engineering, and data center partner teams—contributing to standardized reference designs and building out the HPC engineering function
What We're Looking For
Required experience and capabilities:
- Proven experience in HPC or AI software stack design and delivery at scale—including workload profiling, scheduler configuration (SLURM, PBS, or equivalent), MPI/NCCL tuning, and distributed training frameworks such as PyTorch, JAX, or DeepSpeed
- Deep understanding of GPU software environments: CUDA, cuDNN, NCCL, driver stacks, and the tooling required to run large-scale AI training and inference workloads reliably in production
- Hands-on experience optimizing AI and HPC workloads across multi-GPU and multi-node configurations—including profiling, bottleneck identification, and performance tuning at both application and infrastructure layers
- Strong working knowledge of containerization and orchestration in HPC/AI contexts: Docker, Kubernetes, NVIDIA GPU Operator, and container-native workload management
- Background in an OEM, hyperscaler, or enterprise/research HPC environment with demonstrable exposure to the full design-to-deployment lifecycle for GPU-accelerated workloads
- Ability to produce clear, professional technical documentation and architecture diagrams suitable for both engineering and board-level audiences
- Confident engaging with customers, vendors, and internal engineering teams as a technical authority—able to translate complex software and performance trade-offs into clear, actionable decisions
Nice to have:
- Experience with large-scale cluster performance benchmarking (NCCL tests, MLPerf, or equivalent) and familiarity with performance expectations across different GPU generations and topologies
- Exposure to MLOps tooling and AI platform layers: experiment tracking (MLflow, W&B), model serving frameworks (Triton, vLLM), and pipeline orchestration (Kubeflow, Airflow)
- Familiarity with InfiniBand and high-performance networking as it relates to distributed training performance—sufficient to engage credibly with network and infrastructure teams on topology and tuning decisions
How We Work With You
As a Get A Job.ai candidate, you will apply through our platform at getajob.ai. One of our experienced recruiters will conduct an initial screening to understand your background and career goals. Once we determine mutual fit, we will submit your profile to our client for consideration. Please do not contact the client directly—all communication and coordination will flow through our team to ensure the best experience for everyone involved.
Pay
Compensation details will be discussed during the screening process based on experience and qualifications.
Equal Opportunity
Get A Job.ai is committed to fostering an inclusive recruitment process. We welcome applications from candidates of all backgrounds and work to ensure fair consideration throughout our screening and submission process.
Apply with Get A Job.ai
A recruiter will review your profile and submit you to the client. Do not contact the client directly.
Apply through Get A Job.ai. A recruiter will review your profile and submit you. Do not contact the client directly.
Apply through Get A Job.ai. A recruiter will review your profile and submit you.
Local insights for this role are preparing — this section updates automatically in a few seconds (or refresh).
Listing facts
- Role HPC Cluster Architect
- Employer Get A Job.ai
- Location United Kingdom
- Type Full Time
- Pay (from listing) Pay not listed
- Posted September 5, 2026
- Apply by October 5, 2026
- Country United Kingdom
- Overview Full job description on this page (591 words)
Facts above come from this job record on Get A Job.AI — not copied from third-party review sites.
Limited public data for this employer
We only show facts we can ground in public sources (Wikidata, O*NET, news/discussion links, or this listing). We do not invent Glassdoor-style ratings, salaries, or testimonials when data is thin. Use the listing facts, occupation context, and related openings below while we continue researching.
Explore related openings
Keep exploring on Get A Job.ai
Not quite the right fit? Your next opportunity is a click away.
- Browse all jobs
- More jobs by category
- Remote jobs you can do from anywhere
- Research typical pay for this role
- Set a job alert so new matches reach you first
- Upload your resume to apply faster
Hiring instead? Post a job and reach candidates searching right now.
