Loading...

Senior Site Reliability Engineer — Token Factory (Inference Platform)

  • Company: Get A Job.ai
  • Location: Berlin
  • Salary: Pay not listed
  • Full Time
  • Berlin

Website Get A Job.ai

Represented by Get A Job.ai

About This Opportunity

We are representing a confidential AI cloud infrastructure organization seeking a Senior Site Reliability Engineer to join their inference platform team in Berlin. Our client operates one of the world's largest GPU clouds and is building a next-generation inference platform that serves text, vision, audio, and multimodal foundation models at massive scale.

In this role, you will own the reliability, performance, and observability of the entire inference stack. You'll design and refine telemetry pipelines that transform hundreds of terabytes of signal into actionable insights. Your work will include tuning Kubernetes autoscalers to maximize GPU efficiency, crafting Terraform modules that build resilience into every cluster, and hardening request-routing and retry logic to ensure seamless user experiences even during transient failures.

When incidents occur, you'll leverage the automation and runbooks you've helped create to detect, isolate, and remediate problems within minutes. You'll also drive a post-mortem culture that prevents recurrence and ensures continuous improvement. Your ultimate goal is to scale the platform smoothly while meeting aggressive cost and reliability targets.

Responsibilities

  • Design and maintain telemetry pipelines including metrics, logs, and traces for the inference platform
  • Optimize Kubernetes autoscaling configurations to maximize GPU utilization and efficiency
  • Build and maintain Terraform modules and infrastructure-as-code that embed reliability best practices
  • Harden request-routing, retry logic, and failure-handling mechanisms across distributed systems
  • Create automation, runbooks, and monitoring systems for rapid incident detection and resolution
  • Lead post-mortem processes and implement preventive measures to avoid recurring issues
  • Collaborate with software engineering teams to make reliability a core platform feature
  • Performance-tune the stack from kernel to application layer for high-throughput GPU workloads
  • Design and implement SLOs and alert strategies for production APIs

What We're Looking For

  • Deep expertise with Kubernetes, Prometheus, Grafana, and Terraform in production environments
  • Strong scripting skills in Python or Bash for automation and tooling
  • Proven experience designing alert systems and SLOs for high-throughput distributed APIs
  • Solid understanding of how distributed backends fail in real-world production scenarios
  • Experience with GPU-heavy workloads and frameworks such as vLLM, Triton, or Ray is highly valued
  • Background in MLOps or model-hosting platforms is a strong plus
  • Passion for building self-healing systems and infrastructure automation
  • Strong debugging skills across the full stack from kernel to application
  • Collaborative mindset with ability to work effectively with engineering teams
  • Authorization to work in Germany required

How We Work With You

Candidates apply directly through Get A Job.ai. Our recruiting team will screen your application and conduct an initial interview to understand your background and technical expertise. If there's a strong match, we will submit your profile to our client for consideration. Please note that you should not contact the client directly—all communication will be managed through our talent team to ensure a professional and coordinated process.

Pay

Compensation details will be discussed during the screening process based on experience and qualifications.

Equal Opportunity: Get A Job.ai is committed to inclusive hiring practices and equal employment opportunities for all candidates regardless of race, color, religion, sex, national origin, age, disability, or any other protected characteristic.

Apply with Get A Job.ai

A recruiter will review your profile and submit you to the client. Do not contact the client directly.

More options

Apply with Get A Job.ai

Apply through Get A Job.ai. A recruiter will review your profile and submit you. Do not contact the client directly.

Apply through Get A Job.ai. A recruiter will review your profile and submit you.

Working in Berlin, Deutschland

Weather right now in Berlin, Deutschland: checking… · Local time: · Air quality: · Daylight: · UV index: · Wind: · Pollen:

Berlin is the capital of Germany as well as its largest city by both area and population. With 3.7 million inhabitants, it has the highest population within its city limits of any city in the European Union. The city is also one of the states of Germany, being the third-smallest state in the country by area. Berlin is surrounded by the state of Brandenburg, bordering Brandenburg's capital Potsdam to the southwest. The urban area of Berlin has a population of over 5 million, making it the most populous in Germany. The Berlin-Brandenburg capital region has around 6 million inhabitants and is Ger

Note: Germany observes a public holiday on Oct 3 — German Unity Day.

🇩🇪 Relocation safety for Germany: Very Safevia Warnely, CC BY 4.0

National unemployment rate in Germany: 3.7%via World Bank

GDP per capita in Germany: $60,496via World Bank

Consumer price inflation in Germany: 2.2% (annual) — via World Bank

Real GDP growth in Germany: 0.2% (annual) — via World Bank

Statutory minimum wage in Germany: €2,343/monthvia Eurostat

Cost of living in Germany: 9.1% above the EU averagevia Eurostat

Job vacancy rate in Germany: 2.8%via Eurostat

Average hours worked per year in Germany: 1,332via OECD

Nearby green space: 10 parks within 1.5km — closest is Lustgarten (306m). via OpenStreetMap

Nearest public transit: Staatsoper (bus stop, 55m). via OpenStreetMap

  • Elevation 36m (118 ft)

Job details above are provided by the employer/source. The sections on this page are compiled from public data sources with AI assistance.

Accommodations: if you need a workplace accommodation to apply for or perform this job, see ADA.gov or EEOC.gov for guidance on your rights and how to request one.

Add application deadline to calendar

Listing facts

  • Role Senior Site Reliability Engineer — Token Factory (Inference Platform)
  • Employer Get A Job.ai
  • Location Berlin
  • Type Full Time
  • Pay (from listing) Pay not listed
  • Posted September 15, 2026
  • Apply by October 16, 2026
  • Country Germany
  • Overview Full job description on this page (505 words)

Facts above come from this job record on Get A Job.AI — not copied from third-party review sites.

Limited public data for this employer

We only show facts we can ground in public sources (Wikidata, O*NET, news/discussion links, or this listing). We do not invent Glassdoor-style ratings, salaries, or testimonials when data is thin. Use the listing facts, occupation context, and related openings below while we continue researching.

Explore related openings

Keep exploring on Get A Job.ai

Not quite the right fit? Your next opportunity is a click away.

Hiring instead? Post a job and reach candidates searching right now.

Senior Site Reliability Engineer — Token Factory… Get A Job.ai · Berlin