- Company: Get A Job.ai
- Location: Berlin, Germany
- Salary: Pay not listed
Website Get A Job.ai
Represented by Get A Job.ai
About This Opportunity
Our recruiting team at Get A Job.ai is representing a confidential AI technology company based in Berlin. This organization focuses on combining advanced technology with human expertise to enable data-driven solutions at scale. We're seeking a Senior Site Reliability Engineer to join their platform engineering team and help build reliable, scalable infrastructure that powers machine learning and AI workloads.
This role offers the chance to shape infrastructure strategy across the organization while working with cutting-edge ML and LLM technologies in production environments.
Responsibilities
- Ensure reliability of platform systems including ML and LLM workloads, model serving, and inference infrastructure with GPU-backed endpoints, autoscaling, and performance optimization
- Build and maintain observability systems for ML models, including drift and performance monitoring, LLM-specific tracing, evaluations, and guardrails
- Shape company-wide technical direction and roadmap, establishing golden paths that elevate engineering standards across teams
- Develop developer tooling and automation including reusable GitHub Actions, GitOps workflows, and Terraform modules to accelerate deployment velocity
- Package and maintain reusable components from common open-source tools (Grafana, Istio, CloudNative stack, model registries, feature stores) for multi-environment deployment
- Implement secure-by-default infrastructure with built-in security controls, compliance audits, cost governance, and audit trails
- Manage SLOs, on-call rotation, and incident response covering both infrastructure and ML model performance
What We're Looking For
Required qualifications:
- 5+ years of experience in infrastructure engineering, DevOps, or SRE operating large-scale, high-availability production systems using Kubernetes
- Production operational experience managing live clusters under real load, with proficiency in Helm and Terraform or CloudFormation on at least one major cloud platform (AWS preferred)
- Strong proficiency in Python, Go, or scripting languages for automation and tooling development
- Active use of AI agentic tooling (such as Claude Code or similar) integrated into daily workflows—not just experimentation
- First-principles reasoning ability to evaluate tradeoffs in business terms (reliability vs. velocity, cost vs. blast radius, standardization vs. customization)
- Track record of owning at least one infrastructure build end-to-end with measurable outcomes (deploy time, MTTR, cost, adoption, availability)
- Cross-functional collaboration experience working with product and engineering teams on collective initiatives
- Ability to work across global teams and time zones with strong communication skills
- Willingness to support 24x7 operational processes
Strongly preferred ML and AI platform experience:
- Running ML workloads on Kubernetes including GPU scheduling, capacity planning, and cost management
- Model serving and inference at production scale using tools like KServe, RayServe (preferred), Triton, vLLM, or similar with real latency and cost constraints
- MLOps pipeline tooling experience with training pipelines, model registries, feature stores, and lineage tracking (Kubeflow, MLflow, Feast, Weights & Biases, or equivalents)
- LLMOps in production environments including inference serving, prompt/version management, and LLM observability (tracing, evaluations, drift detection, guardrails, cost per request)
- Governance of ML/LLM workloads as platform capabilities with data-residency controls, PII handling, and audit trails
How We Work With You
When you apply through Get A Job.ai, here's what happens next:
- A recruiter from our talent team will review your application and schedule an initial screening call to discuss your background and interest
- If there's a strong match, we'll coordinate your interview process with our client, providing guidance and support throughout
- We'll submit your profile to the client and manage all communication on your behalf
- Please do not contact the client directly—all coordination happens through Get A Job.ai to ensure the best candidate experience
Pay
Compensation details will be discussed during the screening process based on experience and qualifications.
Equal Opportunity: Get A Job.ai is committed to inclusive recruiting practices. We welcome applications from candidates of all backgrounds and work with our clients to ensure fair and equitable hiring processes.
Apply with Get A Job.ai
A recruiter will review your profile and submit you to the client. Do not contact the client directly.
Apply through Get A Job.ai. A recruiter will review your profile and submit you. Do not contact the client directly.
Apply through Get A Job.ai. A recruiter will review your profile and submit you.
Local insights for this role are preparing — this section updates automatically in a few seconds (or refresh).
Listing facts
- Role Senior Site Reliability Engineer
- Employer Get A Job.ai
- Location Berlin, Germany
- Type Full Time
- Pay (from listing) Pay not listed
- Posted September 3, 2026
- Apply by October 4, 2026
- Country Germany
- Overview Full job description on this page (598 words)
Facts above come from this job record on Get A Job.AI — not copied from third-party review sites.
Typical work in Senior Site Reliability Engineer
Independent occupational context from O*NET (U.S. public-domain labor data). This is about the occupation, not a rewrite of this employer's posting.
- Study product characteristics or customer requirements to determine validation objectives and standards.
- Analyze validation test data to determine whether systems or processes have met validation criteria or to identify root causes of production problems.
- Develop validation master plans, process flow diagrams, test cases, or standard operating procedures.
- Prepare detailed reports or design statements, based on results of validation and qualification tests or reviews of procedures and protocols.
- Maintain validation test equipment.
- Conduct validation or qualification tests of new or existing processes, equipment, or software in accordance with internal protocols or external standards.
Limited public data for this employer
We only show facts we can ground in public sources (Wikidata, O*NET, news/discussion links, or this listing). We do not invent Glassdoor-style ratings, salaries, or testimonials when data is thin. Use the listing facts, occupation context, and related openings below while we continue researching.
Explore related openings
Occupation family: Senior Site Reliability Engineer
Keep exploring on Get A Job.ai
Not quite the right fit? Your next opportunity is a click away.
- Browse all jobs
- More jobs by category
- Remote jobs you can do from anywhere
- Research typical pay for this role
- Set a job alert so new matches reach you first
- Upload your resume to apply faster
Hiring instead? Post a job and reach candidates searching right now.
