- Company: Get A Job.ai
- Location: Sydney, Australia
- Salary: Pay not listed
Website Get A Job.ai
Represented by Get A Job.ai
About This Opportunity
We are representing a confidential AI infrastructure organization operating in the Asia Pacific region. Our client is seeking a Senior AI Infrastructure Engineer with deep Kubernetes expertise to own the technical design and delivery of their GPU-accelerated platform. This is a hands-on principal-level individual contributor role based in Sydney, Australia.
You will be responsible for building production-grade cluster lifecycle, control-plane, networking, storage, security, observability, and automation capabilities across bare-metal environments. This role requires solving the hardest platform engineering problems and setting engineering standards for a secure, resilient, multi-tenant platform deployed at scale.
Responsibilities
- Define and own the Kubernetes platform reference architecture across management and workload clusters, including control-plane topology, cluster lifecycle, multi-tenancy, workload isolation, and failure-domain design
- Build and maintain backend services, APIs, controllers, operators, and automation required to provision, configure, upgrade, scale, and retire Kubernetes clusters reliably
- Engineer repeatable bare-metal Kubernetes deployment and lifecycle workflows using infrastructure-as-code and automated provisioning technologies such as Cluster API, kubeadm, Redfish, PXE, Ironic, or Metal3
- Design and operate cluster networking across CNI, ingress, service discovery, DNS, load balancing, network policy, and service mesh; integrate Multus, SR-IOV, BGP, InfiniBand, or RoCE for high-performance AI workloads
- Define persistent-storage and data-service patterns using CSI, Ceph, local NVMe, object storage, backup and restore, and disaster-recovery mechanisms for stateful platform and AI workloads
- Integrate and productionise NVIDIA GPU and Network Operators, device plugins, drivers, DCGM telemetry, scheduling, quotas, and topology-aware placement for multi-node accelerated workloads
- Establish GitOps and CI/CD patterns for platform software, configuration, policy, and release management with safe testing, progressive rollout, rollback, and upgrade practices
- Build platform security through identity and access control, RBAC, secrets management, policy-as-code, image and software-supply-chain controls, tenant isolation, and auditable change management
- Define service-level objectives and engineer observability for metrics, logs, traces, events, capacity, and performance; lead diagnosis of complex distributed systems failures
- Set engineering standards, design patterns, review practices, and operational readiness criteria; mentor senior engineers and resolve cross-team technical decisions
What We're Looking For
- 7+ years of progressive infrastructure, systems, or platform engineering experience, including substantial ownership of production Kubernetes platforms
- At least 3 years operating at senior staff, principal, or equivalent level
- Deep knowledge of Kubernetes internals, including API server, etcd, scheduler, controller manager, kubelet, admission, CRI, CNI, CSI, reconciliation patterns, cluster performance, upgrades, and control-plane failure modes
- Demonstrated experience designing, building, and operating highly available, large-scale, multi-cluster Kubernetes platforms on bare metal, private cloud, or hybrid infrastructure
- Strong software engineering ability in Go and/or Rust, with practical Python and Bash skills; experience building Kubernetes operators, controllers, admission webhooks, CLIs, or platform services
- Expert Linux systems knowledge, including namespaces, cgroups, systemd, kernel, host networking and container runtime behaviour, performance analysis, and low-level troubleshooting
- Strong Kubernetes networking expertise across Cilium, Calico, or equivalent CNI implementations, plus load balancing, DNS, ingress, BGP, network policy, and multi-network architectures
- Strong infrastructure automation and GitOps experience with tools such as Terraform, Ansible, Argo CD, Flux, GitHub Actions, GitLab CI, or Jenkins
- Practical experience with Kubernetes security and governance, including RBAC, OPA Gatekeeper or Kyverno, secrets management, certificate lifecycle, image security, and workload isolation
- Experience implementing production observability with Prometheus, Grafana, OpenTelemetry, Loki, Elasticsearch, or equivalent technologies
- Experience with GPU-enabled Kubernetes infrastructure, NVIDIA GPU Operator, accelerator scheduling for AI workloads at large scale, RDMA networking, and distributed AI workload requirements
- Experience with distributed storage and data services such as Ceph, CSI-backed storage, object storage, backup and restore, and disaster recovery
- CKA-level expertise expected; CKA, CKS, or relevant cloud-native certifications strongly preferred
- Bachelor's degree in computer science, engineering, or related discipline, or equivalent depth of practical engineering experience
- Clear technical judgement and communication, with a record of influencing architecture across software, networking, security, platform, and operations teams
How We Work With You
Candidates apply directly through Get A Job.ai. Our recruiting team will screen your application and conduct an initial interview. If there is a strong match, we will submit your profile to our client for consideration. Please do not contact the client directly, as all communication is managed through our talent team to ensure a professional and confidential process.
Pay
Compensation details will be discussed during the screening process based on experience and qualifications.
Equal Opportunity: Get A Job.ai is committed to inclusive hiring practices. We welcome applications from candidates of all backgrounds and work with clients who share our commitment to building diverse teams.
Apply with Get A Job.ai
A recruiter will review your profile and submit you to the client. Do not contact the client directly.
Apply through Get A Job.ai. A recruiter will review your profile and submit you. Do not contact the client directly.
Apply through Get A Job.ai. A recruiter will review your profile and submit you.
Terms used in this posting
- hybrid
- A work arrangement combining both in-office and remote/at-home work, typically on a set schedule.
Explore Get A Job.ai online
Working in Sydney, Sydney Region
Weather right now in Sydney, Sydney Region: checking… · Local time: · Air quality: · Daylight: · UV index: · Wind: · Pollen:
Sydney is the capital city of the state of New South Wales, and the most populous city in Australia. Located on Australia's east coast, the metropolis surrounds Sydney Harbour and extends about 80 kilometres (50 mi) from the Pacific Ocean in the east to the Blue Mountains in the west, and about 80 kilometres (50 mi) from Ku-ring-gai Chase National Park and the Hawkesbury River in the north and north-west, to the Royal National Park and Macarthur in the south and south-west. Greater Sydney consists of 658 suburbs, spread across 33 local government areas. Residents of the city are colloquially k
New South Wales is a state on the east coast of Australia. It borders Queensland to the north, Victoria to the south, and South Australia to the west. Its coast borders the Coral and Tasman Seas to the east. The Australian Capital Territory and Jervis Bay Territory are enclaves within the state. New South Wales' state capital is Sydney, which is also Australia's most populous city. As of September
🇦🇺 Relocation safety for Australia: Very Safe — via Warnely, CC BY 4.0
National unemployment rate in Australia: 4.1% — via World Bank
GDP per capita in Australia: $65,130 — via World Bank
Consumer price inflation in Australia: 2.9% (annual) — via World Bank
Real GDP growth in Australia: 1.4% (annual) — via World Bank
Average hours worked per year in Australia: 1,633 — via OECD
Recent seismic activity: 1 earthquake (M4.5+) within 200km in the last 6 months — largest M4.5 near 24 km ENE of Canowindra, Australia. via USGS
Nearby green space: 15 parks within 1.5km — closest is Charles Throsby reserve (508m). via OpenStreetMap
Nearest public transit: Anderson Rd opp Anzac Ave (bus stop, 736m). via OpenStreetMap
Average weekly earnings in New South Wales: A$1,616 — via Australian Bureau of Statistics
Unemployment rate in New South Wales: 4.1% — via Australian Bureau of Statistics
- Elevation 88m (289 ft)
Source: Wikipedia (state)
Job details above are provided by the employer/source. The sections on this page are compiled from public data sources with AI assistance.
Accommodations: if you need a workplace accommodation to apply for or perform this job, see ADA.gov or EEOC.gov for guidance on your rights and how to request one.
Listing facts
- Role Senior AI Infrastructure Engineer, Kubernetes
- Employer Get A Job.ai
- Location Sydney, Australia
- Type Full Time
- Pay (from listing) Pay not listed
- Posted September 13, 2026
- Apply by October 13, 2026
- Country Australia
- Overview Full job description on this page (726 words)
Facts above come from this job record on Get A Job.AI — not copied from third-party review sites.
Limited public data for this employer
We only show facts we can ground in public sources (Wikidata, O*NET, news/discussion links, or this listing). We do not invent Glassdoor-style ratings, salaries, or testimonials when data is thin. Use the listing facts, occupation context, and related openings below while we continue researching.
Explore related openings
Keep exploring on Get A Job.ai
Not quite the right fit? Your next opportunity is a click away.
- Browse all jobs
- More jobs by category
- Remote jobs you can do from anywhere
- Research typical pay for this role
- Set a job alert so new matches reach you first
- Upload your resume to apply faster
Hiring instead? Post a job and reach candidates searching right now.
