Loading...

Technical Lead – GPU Infrastructure

  • Company: Get A Job.ai
  • Location: Remote job
  • Salary: Pay not listed
  • Work type: Remote

Website Get A Job.ai

Represented by Get A Job.ai

About the Opportunity

We are representing a confidential digital finance organization in their search for a Technical Lead - GPU Infrastructure to own and scale their GPU compute and managed inference platform. This is a hands-on infrastructure leadership role with a fixed delivery window in its first six months—you'll architect and deliver the full stack on bare-metal GPU infrastructure while leading a distributed engineering team of approximately twelve people across backend, frontend, DevOps, QA, and documentation.

This platform orchestrates GPU workloads for research, model training, and inference at scale. You'll be building a managed Slurm scheduling layer for internal teams and a custom Kubernetes control plane for inference tenancy—transitioning the platform from managed cluster orchestration to full-stack ownership on bare metal.

You will report to the Senior Technical Product Manager and serve as the primary technical interface to infrastructure partners, translating requirements into specifications and acceptance tests while managing delivery, architecture, and team leadership.

Responsibilities

  • Architecture ownership: Drive architecture proposals, high-level and low-level designs through review, and maintain them as the baseline for platform evolution
  • Team leadership: Line-manage a distributed team across Node.js backend, React frontend, DevOps, QA, and documentation—setting engineering standards, conducting code and design reviews, managing release gates, one-to-ones, and performance input
  • Bare-metal GPU scheduling: Design, build, and operate a managed Slurm service including controller and accounting, partitions and login nodes, node onboarding and acceptance, driver and CUDA lifecycle, stalled-job and health detection, drain and autohealing, and storage visibility
  • Kubernetes control plane: Own cluster bootstrap and lifecycle on partner-provided bare metal, NVIDIA GPU Operator and Network Operator, VM-based GPU isolation using KubeVirt and VFIO, plus day-2 operations including upgrades, backup and recovery, and node replacement
  • Managed inference at scale: Design serving architecture with multi-GPU and multi-node parallelism, autoscaling, request routing, endpoint reliability, and confidential-compute-capable capacity
  • Observability and operations: Implement metrics, logging, alerting, and SLOs across control plane, GPU fleet, and application tiers; lead incident response and post-incident review; establish a sustainable on-call model
  • Partner and vendor interface: Serve as primary technical contact to infrastructure partners and vendors—turn requirements into written specifications and acceptance tests, run escalations to closure, and contribute to capacity planning and hardware sourcing
  • Internal stakeholder management: Work directly with research, model-training, and product teams to translate workloads into platform requirements and broker capacity when constrained
  • Team building: Complete the platform team and set the technical bar for future hires

What We're Looking For

Must have:

  • Eight or more years of hands-on engineering experience, including at least three years leading teams that build and operate infrastructure platforms depended upon by other teams
  • Bachelor's or Master's degree in computer science or engineering, or equivalent practical experience
  • Slurm at scale, hands-on: Real experience running slurmctld and slurmdbd for production users—partitions, QoS and priority, accounting, prolog and epilog, node health scripting, and upgrades with jobs running. Ideally operated an HPC or GPU training cluster for a research population
  • GPU fleet operation on bare metal: NVIDIA driver and CUDA lifecycle management, Fabric Manager and NVSwitch behaviour on SXM systems, DCGM-based health and utilisation, MIG, node burn-in and acceptance
  • High-performance interconnects: InfiniBand fabric and subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance issues
  • Linux systems depth: Kernel modules and drivers, PCIe passthrough and vfio-pci, cgroups and namespaces, performance tuning for compute-heavy workloads
  • Production Kubernetes operation: Control plane management, upgrades, CNI and CSI, operators and custom controllers, multi-tenancy design—not just deployment experience
  • HPC storage and data movement: Shared filesystems (VAST, Lustre, NFS), node-local NVMe caching, distributing large model weights and datasets across many nodes
  • Observability and operations: Prometheus, Grafana, Loki or equivalents, SLOs, incident response, and post-incident review
  • Working fluency in JavaScript and Node.js sufficient to review a control plane, CLI, and worker services with authority and make architecture decisions (not a feature-development requirement)
  • A shipped platform with real users—multi-tenant IaaS or PaaS, or a research computing service including resource isolation, quotas, usage metering, and user-facing API and CLI surfaces
  • Leadership that stays in the code: people management across time zones, cross-track review, written architecture decisions with alternatives recorded, and the ability to tell a partner or executive no with documented reasons
  • Excellent written and spoken English—most partner and leadership work happens in writing
  • Location: Fully remote, based between UTC and UTC+5:30 so your working day overlaps both Europe and India where the team and partners work. Occasional travel to partner sites and team events required

Desirable:

  • Slurm operators on Kubernetes (Soperator, Slinky) or Kubernetes-native schedulers (Kueue, Volcano, KAI, Kubeflow Trainer)
  • Modern serving stacks (vLLM, SGLang, TensorRT-LLM): parallelism strategies, quantisation trade-offs, GPU memory planning
  • VM and container isolation for multi-tenant GPU compute (KubeVirt, Kata Containers, QEMU and KVM, Firecracker); confidential computing (Intel TDX, AMD SEV-SNP, NVIDIA confidential-compute mode)
  • Cluster API and kubeadm, Cilium, NVSentinel-class autohealing, infrastructure as code and GitOps
  • Time on the operator side of a GPU cloud, a national or university HPC centre, or an AI lab's platform team
  • Peer-to-peer or distributed-systems background
  • Experience with a hardware provider who provisions but does not operate, and turning that relationship into a written contract with acceptance tests

How We Work with You

Candidates apply through Get A Job.ai. A recruiter from our talent team will screen your application and conduct an initial conversation. If there's a strong match, we will submit your profile to our client for consideration. Please do not contact the client directly—all communication and coordination will flow through Get A Job.ai to ensure a smooth and professional process for everyone involved.

Pay

Compensation details will be discussed during the screening process and are competitive for this level of technical leadership in a remote, global organization.

Equal Opportunity: Get A Job.ai is committed to inclusive recruiting practices. We welcome applications from candidates of all backgrounds and work to ensure every applicant receives fair consideration.

Apply with Get A Job.ai

A recruiter will review your profile and submit you to the client. Do not contact the client directly.

More options

Apply with Get A Job.ai

Apply through Get A Job.ai. A recruiter will review your profile and submit you. Do not contact the client directly.

Apply through Get A Job.ai. A recruiter will review your profile and submit you.

Terms used in this posting

on-call
You may be required to be reachable and available to work outside normal scheduled hours, typically for a set rotation.

Working in Remote job

    Job details above are provided by the employer/source. The sections on this page are compiled from public data sources with AI assistance.

    Accommodations: if you need a workplace accommodation to apply for or perform this job, see ADA.gov or EEOC.gov for guidance on your rights and how to request one.

    Add application deadline to calendar

    Listing facts

    • Role Technical Lead – GPU Infrastructure
    • Employer Get A Job.ai
    • Location Remote job · Remote-friendly
    • Type Full Time
    • Pay (from listing) Pay not listed
    • Posted September 17, 2026
    • Apply by October 17, 2026
    • Overview Full job description on this page (969 words)

    Facts above come from this job record on Get A Job.AI — not copied from third-party review sites.

    Limited public data for this employer

    We only show facts we can ground in public sources (Wikidata, O*NET, news/discussion links, or this listing). We do not invent Glassdoor-style ratings, salaries, or testimonials when data is thin. Use the listing facts, occupation context, and related openings below while we continue researching.

    Explore related openings

    Keep exploring on Get A Job.ai

    Not quite the right fit? Your next opportunity is a click away.

    Hiring instead? Post a job and reach candidates searching right now.

    Technical Lead – GPU Infrastructure Get A Job.ai · Remote job