Loading...

AI Engineer – Inference

  • Company: Get A Job.ai
  • Location: Sydney, Australia
  • Salary: Pay not listed

Website Get A Job.ai

Represented by Get A Job.ai

Responsibilities

We are representing a confidential AI infrastructure organization seeking an AI Engineer specializing in inference systems. In this role, you will build and improve production-scale inference capabilities, making models available as reliable, secure, and high-performance endpoints for internal products, external customers, and future service offerings.

You will establish the engineering foundation for self-hosted model serving in a GPU cloud environment. This includes model onboarding, deployment, endpoint provisioning, runtime selection, performance benchmarking and optimization, observability, capacity management, security, and operational lifecycle management.

Your key responsibilities include:

  • Build, operate, and continuously improve self-hosted AI inference services for internal applications, customer-facing products, and future inference-as-a-service offerings
  • Define and implement standard model-onboarding workflows covering intake, validation, packaging, runtime selection, optimization, deployment, and lifecycle management
  • Provision and manage secure, scalable inference endpoints for interactive generation, RAG, embeddings, reranking, batch processing, multimodal use cases, and agentic workflows
  • Develop reusable deployment templates, APIs, SDKs, configuration standards, and self-service workflows for endpoint management
  • Work with leading inference frameworks such as TensorRT-LLM, SGLang, vLLM, Triton Inference Server, and related serving tools
  • Optimize model-serving performance using quantization, compilation, batching, request routing, KV-cache management, speculative decoding, and distributed parallelism
  • Build and validate reusable inference recipes specifying model versions, framework configurations, GPU requirements, and expected performance envelopes
  • Apply optimization approaches including NVFP4, FP8, INT8, TensorRT compilation, and efficient attention mechanisms while maintaining quality targets
  • Design distributed inference configurations for large models using tensor, pipeline, expert, and data parallelism
  • Define endpoint resource profiles, placement requirements, autoscaling rules, and workload-management policies
  • Build benchmarking and qualification workflows using controlled experiments, load tests, and performance profiling
  • Measure and improve key metrics including time-to-first-token, inter-token latency, throughput, GPU utilization, and cost efficiency
  • Establish automated performance-regression testing for model versions, runtime upgrades, and infrastructure changes
  • Build operational observability for endpoint availability, request volume, latency, errors, and capacity
  • Partner with applications teams to provide endpoints for agent planning, retrieval, tool use, and autonomous workflows
  • Collaborate with Product, DevOps, Platform, Infrastructure, and Security teams to ensure clear and secure inference provisioning

What We're Looking For

Required Experience:

  • 5+ years of software engineering experience, including 3+ years in AI inference, model serving, ML systems, high-performance computing, or distributed systems
  • Demonstrated experience building or operating production model-serving platforms, inference APIs, GPU-backed services, or multi-tenant AI systems
  • Hands-on experience with modern inference frameworks such as TensorRT-LLM, TensorRT, SGLang, vLLM, Triton Inference Server, or equivalent technologies
  • Strong understanding of the NVIDIA AI software stack including CUDA, cuDNN, NCCL, TensorRT, and GPU profiling
  • Practical understanding of LLM and generative-AI serving behavior including batching, context length, concurrency, KV-cache management, and latency-throughput trade-offs
  • Experience with model optimization methods including quantization, compilation, mixed precision, kernel fusion, and parallelism
  • Strong Python skills and working proficiency in C++ or Go for inference services, APIs, and performance-critical development
  • Experience with distributed inference patterns including tensor, pipeline, expert, and data parallelism
  • Familiarity with Kubernetes, containers, CI/CD, service APIs, autoscaling, and production platform operations
  • Understanding of high-performance GPU infrastructure including GPU topology, NVLink, NVSwitch, RDMA, and network fabrics
  • Experience with inference benchmarking, performance profiling, load testing, and analysis of throughput, latency, and utilization metrics
  • Familiarity with model-serving use cases such as RAG, embeddings, multimodal inference, and agentic applications
  • Understanding of security and governance for inference services including authentication, authorization, tenant isolation, and audit logging

Key Competencies:

  • Self-hosted inference platform engineering and production ownership
  • High-performance LLM and generative-AI optimization
  • Development of validated, repeatable, versioned inference recipes and deployment configurations
  • Quantitative performance engineering across latency, throughput, GPU utilization, memory, scaling, and cost
  • Cross-functional collaboration with applications, DevOps, platform, and operations teams

Location: Sydney, Australia

Employment Type: Permanent full-time

How We Work With You

When you apply through Get A Job.ai, our recruiting team will review your qualifications and conduct an initial screening. If there's a strong match, we'll submit your profile to our client for consideration. Throughout the process, we'll keep you informed and provide guidance.

Please do not contact the client directly. All applications and inquiries should go through Get A Job.ai to ensure proper coordination and the best candidate experience.

Pay

Compensation details will be discussed during the screening process based on experience and qualifications.

Equal Opportunity: Get A Job.ai is committed to inclusive recruiting practices. We welcome applications from candidates of all backgrounds and work to ensure fair consideration throughout the hiring process.

Apply with Get A Job.ai

A recruiter will review your profile and submit you to the client. Do not contact the client directly.

More options

Apply with Get A Job.ai

Apply through Get A Job.ai. A recruiter will review your profile and submit you. Do not contact the client directly.

Apply through Get A Job.ai. A recruiter will review your profile and submit you.

Working in Sydney, Sydney Region

Weather right now in Sydney, Sydney Region: checking… · Local time: · Air quality: · Daylight: · UV index: · Wind: · Pollen:

Sydney is the capital city of the state of New South Wales, and the most populous city in Australia. Located on Australia's east coast, the metropolis surrounds Sydney Harbour and extends about 80 kilometres (50 mi) from the Pacific Ocean in the east to the Blue Mountains in the west, and about 80 kilometres (50 mi) from Ku-ring-gai Chase National Park and the Hawkesbury River in the north and north-west, to the Royal National Park and Macarthur in the south and south-west. Greater Sydney consists of 658 suburbs, spread across 33 local government areas. Residents of the city are colloquially k

New South Wales is a state on the east coast of Australia. It borders Queensland to the north, Victoria to the south, and South Australia to the west. Its coast borders the Coral and Tasman Seas to the east. The Australian Capital Territory and Jervis Bay Territory are enclaves within the state. New South Wales' state capital is Sydney, which is also Australia's most populous city. As of September

🇦🇺 Relocation safety for Australia: Very Safevia Warnely, CC BY 4.0

National unemployment rate in Australia: 4.1%via World Bank

GDP per capita in Australia: $65,130via World Bank

Consumer price inflation in Australia: 2.9% (annual) — via World Bank

Real GDP growth in Australia: 1.4% (annual) — via World Bank

Average hours worked per year in Australia: 1,633via OECD

Recent seismic activity: 1 earthquake (M4.5+) within 200km in the last 6 months — largest M4.5 near 24 km ENE of Canowindra, Australia. via USGS

Nearby green space: 15 parks within 1.5km — closest is Charles Throsby reserve (508m). via OpenStreetMap

Nearest public transit: Anderson Rd opp Anzac Ave (bus stop, 736m). via OpenStreetMap

Average weekly earnings in New South Wales: A$1,616via Australian Bureau of Statistics

Unemployment rate in New South Wales: 4.1%via Australian Bureau of Statistics

  • Elevation 88m (289 ft)

Source: Wikipedia (state)

Job details above are provided by the employer/source. The sections on this page are compiled from public data sources with AI assistance.

Accommodations: if you need a workplace accommodation to apply for or perform this job, see ADA.gov or EEOC.gov for guidance on your rights and how to request one.

Add application deadline to calendar

Listing facts

  • Role AI Engineer – Inference
  • Employer Get A Job.ai
  • Location Sydney, Australia
  • Type Full Time
  • Pay (from listing) Pay not listed
  • Posted September 13, 2026
  • Apply by October 13, 2026
  • Country Australia
  • Overview Full job description on this page (714 words)

Facts above come from this job record on Get A Job.AI — not copied from third-party review sites.

Limited public data for this employer

We only show facts we can ground in public sources (Wikidata, O*NET, news/discussion links, or this listing). We do not invent Glassdoor-style ratings, salaries, or testimonials when data is thin. Use the listing facts, occupation context, and related openings below while we continue researching.

Explore related openings

Keep exploring on Get A Job.ai

Not quite the right fit? Your next opportunity is a click away.

Hiring instead? Post a job and reach candidates searching right now.

AI Engineer – Inference Get A Job.ai · Sydney, Australia