- Company: Get A Job.ai
- Location: Berlin
- Salary: Pay not listed
Website Get A Job.ai
Represented by Get A Job.ai
About This Opportunity
We are representing a confidential AI cloud infrastructure organization seeking a Senior Site Reliability Engineer to join their inference platform team in Berlin. Our client operates one of the world's largest GPU clouds and is building a next-generation inference platform that serves text, vision, audio, and multimodal foundation models at massive scale.
In this role, you will own the reliability, performance, and observability of the entire inference stack. You'll design and refine telemetry pipelines that transform hundreds of terabytes of signal into actionable insights. Your work will include tuning Kubernetes autoscalers to maximize GPU efficiency, crafting Terraform modules that build resilience into every cluster, and hardening request-routing and retry logic to ensure seamless user experiences even during transient failures.
When incidents occur, you'll leverage the automation and runbooks you've helped create to detect, isolate, and remediate problems within minutes. You'll also drive a post-mortem culture that prevents recurrence and ensures continuous improvement. Your ultimate goal is to scale the platform smoothly while meeting aggressive cost and reliability targets.
Responsibilities
- Design and maintain telemetry pipelines including metrics, logs, and traces for the inference platform
- Optimize Kubernetes autoscaling configurations to maximize GPU utilization and efficiency
- Build and maintain Terraform modules and infrastructure-as-code that embed reliability best practices
- Harden request-routing, retry logic, and failure-handling mechanisms across distributed systems
- Create automation, runbooks, and monitoring systems for rapid incident detection and resolution
- Lead post-mortem processes and implement preventive measures to avoid recurring issues
- Collaborate with software engineering teams to make reliability a core platform feature
- Performance-tune the stack from kernel to application layer for high-throughput GPU workloads
- Design and implement SLOs and alert strategies for production APIs
What We're Looking For
- Deep expertise with Kubernetes, Prometheus, Grafana, and Terraform in production environments
- Strong scripting skills in Python or Bash for automation and tooling
- Proven experience designing alert systems and SLOs for high-throughput distributed APIs
- Solid understanding of how distributed backends fail in real-world production scenarios
- Experience with GPU-heavy workloads and frameworks such as vLLM, Triton, or Ray is highly valued
- Background in MLOps or model-hosting platforms is a strong plus
- Passion for building self-healing systems and infrastructure automation
- Strong debugging skills across the full stack from kernel to application
- Collaborative mindset with ability to work effectively with engineering teams
- Authorization to work in Germany required
How We Work With You
Candidates apply directly through Get A Job.ai. Our recruiting team will screen your application and conduct an initial interview to understand your background and technical expertise. If there's a strong match, we will submit your profile to our client for consideration. Please note that you should not contact the client directly—all communication will be managed through our talent team to ensure a professional and coordinated process.
Pay
Compensation details will be discussed during the screening process based on experience and qualifications.
Equal Opportunity: Get A Job.ai is committed to inclusive hiring practices and equal employment opportunities for all candidates regardless of race, color, religion, sex, national origin, age, disability, or any other protected characteristic.
Apply with Get A Job.ai
A recruiter will review your profile and submit you to the client. Do not contact the client directly.
Apply through Get A Job.ai. A recruiter will review your profile and submit you. Do not contact the client directly.
Apply through Get A Job.ai. A recruiter will review your profile and submit you.
Local insights for this role are preparing — this section updates automatically in a few seconds (or refresh).
Listing facts
- Role Senior Site Reliability Engineer — Token Factory (Inference Platform)
- Employer Get A Job.ai
- Location Berlin
- Type Full Time
- Pay (from listing) Pay not listed
- Posted September 15, 2026
- Apply by October 16, 2026
- Country Germany
- Overview Full job description on this page (505 words)
Facts above come from this job record on Get A Job.AI — not copied from third-party review sites.
Limited public data for this employer
We only show facts we can ground in public sources (Wikidata, O*NET, news/discussion links, or this listing). We do not invent Glassdoor-style ratings, salaries, or testimonials when data is thin. Use the listing facts, occupation context, and related openings below while we continue researching.
Explore related openings
Keep exploring on Get A Job.ai
Not quite the right fit? Your next opportunity is a click away.
- Browse all jobs
- More jobs by category
- Remote jobs you can do from anywhere
- Research typical pay for this role
- Set a job alert so new matches reach you first
- Upload your resume to apply faster
Hiring instead? Post a job and reach candidates searching right now.
