Loading...

Senior AI Infrastructure Engineer, Kubernetes

  • Company: Get A Job.ai
  • Location: Sydney, Australia
  • Salary: Pay not listed

Website Get A Job.ai

Represented by Get A Job.ai

About This Opportunity

We are representing a confidential AI infrastructure organization in Sydney, Australia, seeking a Senior AI Infrastructure Engineer with deep Kubernetes expertise. This is a hands-on principal-level individual contributor role where you will own the technical design and delivery of the backend infrastructure that powers a production GPU cloud platform.

You will solve the hardest platform engineering problems, set Kubernetes engineering standards, and provide domain-level technical sign-off for platform designs. Working across AI platforms, solutions architecture, networking, security, and operations teams, you will create a secure, resilient, multi-tenant platform deployed at scale across GPU-accelerated bare-metal environments.

Responsibilities

  • Define and own the Kubernetes platform reference architecture across management and workload clusters, including control-plane topology, cluster lifecycle, multi-tenancy, workload isolation, and failure-domain design
  • Build and maintain the backend services, APIs, controllers, operators, and automation required to provision, configure, upgrade, scale, and retire Kubernetes clusters reliably
  • Engineer repeatable bare-metal Kubernetes deployment and lifecycle workflows using infrastructure-as-code and automated provisioning technologies such as Cluster API, kubeadm, Redfish, PXE, Ironic, or Metal3
  • Design and operate cluster networking across CNI, ingress, service discovery, DNS, load balancing, network policy, and service mesh; integrate Multus, SR-IOV, BGP, InfiniBand, or RoCE where required for high-performance AI workloads
  • Define persistent-storage and data-service patterns using CSI, Ceph, local NVMe, object storage, backup and restore, and disaster-recovery mechanisms appropriate for stateful platform and AI workloads
  • Integrate and productionize NVIDIA GPU and Network Operators, device plugins, drivers, DCGM telemetry, scheduling, quotas, and topology-aware placement for multi-node accelerated workloads
  • Establish GitOps and CI/CD patterns for platform software, configuration, policy, and release management, with safe testing, progressive rollout, rollback, and upgrade practices
  • Build platform security into the architecture through identity and access control, RBAC, secrets management, policy-as-code, image and software-supply-chain controls, tenant isolation, and auditable change management
  • Define service-level objectives and engineer observability for metrics, logs, traces, events, capacity, and performance; lead diagnosis of complex distributed systems failures and eliminate recurring operational toil
  • Set engineering standards, design patterns, review practices, and operational readiness criteria; mentor senior engineers and resolve cross-team technical decisions while remaining directly involved in implementation

What We're Looking For

  • 7+ years of progressive infrastructure, systems, or platform engineering experience, including substantial ownership of production Kubernetes platforms and at least 3 years operating at senior staff, principal, or equivalent level
  • Deep knowledge of Kubernetes internals, including the API server, etcd, scheduler, controller manager, kubelet, admission, CRI, CNI, CSI, reconciliation patterns, cluster performance, upgrades, and control-plane failure modes
  • Demonstrated experience designing, building, and operating highly available, large scale and multi-cluster Kubernetes platforms on bare metal, private cloud, or hybrid infrastructure
  • Strong software engineering ability in Go and/or Rust, with practical Python and Bash skills; experience building Kubernetes operators, controllers, admission webhooks, CLIs, or platform services
  • Expert Linux systems knowledge, including namespaces, cgroups, systemd, kernel, host networking and container runtime behavior, performance analysis, and low-level troubleshooting
  • Strong Kubernetes networking expertise across Cilium, Calico, or equivalent CNI implementations, plus load balancing, DNS, ingress, BGP, network policy, and multi-network architectures
  • Strong infrastructure automation and GitOps experience with tools such as Terraform, Ansible, Argo CD, Flux, GitHub Actions, GitLab CI, or Jenkins
  • Practical experience with Kubernetes security and governance, including RBAC, OPA Gatekeeper or Kyverno, secrets management, certificate lifecycle, image security, and workload isolation
  • Experience implementing production observability with Prometheus, Grafana, OpenTelemetry, Loki, Elasticsearch, or equivalent technologies, and using telemetry to manage reliability, capacity, and performance
  • Experience with GPU-enabled Kubernetes infrastructure, NVIDIA GPU Operator, accelerator scheduling for AI workloads at large scale, RDMA networking, and distributed AI workload requirements
  • Experience with distributed storage and data services such as Ceph, CSI-backed storage, object storage, backup and restore, and disaster recovery
  • CKA-level expertise is expected; CKA, CKS, or relevant cloud-native certifications are strongly preferred
  • Bachelor's degree in computer science, engineering, or a related discipline, or equivalent depth of practical engineering experience
  • Clear technical judgment and communication, with a record of influencing architecture across software, networking, security, platform, and operations teams

Location: Sydney, Australia (full-time)

How We Work With You

Our talent team at Get A Job.ai represents exceptional candidates to leading organizations. When you apply through our platform, a dedicated recruiter will review your profile and contact you to discuss this opportunity in detail. If there's mutual interest, we will submit your application to our client for consideration.

Please do not contact the client organization directly. All applications and communications should flow through Get A Job.ai to ensure a professional and coordinated process.

Pay

Compensation details will be discussed during the screening process with our recruiting team.

Equal Employment Opportunity: Get A Job.ai is committed to inclusive recruiting practices and welcomes applications from candidates of all backgrounds. We evaluate all qualified applicants without regard to race, color, religion, sex, national origin, age, disability, veteran status, or any other protected characteristic.

Apply with Get A Job.ai

A recruiter will review your profile and submit you to the client. Do not contact the client directly.

More options

Apply with Get A Job.ai

Apply through Get A Job.ai. A recruiter will review your profile and submit you. Do not contact the client directly.

Apply through Get A Job.ai. A recruiter will review your profile and submit you.

Local insights for this role are preparing — this section updates automatically in a few seconds (or refresh).

Listing facts

  • Role Senior AI Infrastructure Engineer, Kubernetes
  • Employer Get A Job.ai
  • Location Sydney, Australia
  • Type Full Time
  • Pay (from listing) Pay not listed
  • Posted September 13, 2026
  • Apply by October 13, 2026
  • Country Australia
  • Overview Full job description on this page (795 words)

Facts above come from this job record on Get A Job.AI — not copied from third-party review sites.

Limited public data for this employer

We only show facts we can ground in public sources (Wikidata, O*NET, news/discussion links, or this listing). We do not invent Glassdoor-style ratings, salaries, or testimonials when data is thin. Use the listing facts, occupation context, and related openings below while we continue researching.

Explore related openings

Keep exploring on Get A Job.ai

Not quite the right fit? Your next opportunity is a click away.

Hiring instead? Post a job and reach candidates searching right now.

Senior AI Infrastructure Engineer, Kubernetes Get A Job.ai · Sydney, Australia