Loading...

Senior Site Reliability Engineer / SRE – Kubernetes & Hybrid Cloud (m/f/d)

  • Company: Get A Job.ai
  • Location: Berlin
  • Salary: Pay not listed
  • Full Time
  • Berlin

Website Get A Job.ai

Represented by Get A Job.ai

About This Opportunity

We are representing a confidential e-commerce technology company operating throughout Europe. They are building a modern private cloud platform on Kubernetes and their own infrastructure—this is genuine infrastructure engineering, not managed cloud administration. Our talent team is seeking a Senior Site Reliability Engineer to help define SRE practices from the ground up alongside their Berlin development teams.

This role offers the rare opportunity to work with on-premises Kubernetes clusters running on physical hardware in Frankfurt, with plans for hybrid cloud capabilities. You will be joining as the platform scales to handle billions of queries annually for major European retailers. A Team Lead SRE position is being hired concurrently, so you will help build the practice rather than inherit established processes.

Responsibilities

  • Define and own service level objectives, indicators, and error budgets to drive data-informed reliability decisions
  • Lead end-to-end incident response including detection, communication, blameless postmortems, and structural improvements to prevent entire classes of incidents
  • Eliminate operational toil through automation and GitOps practices
  • Evolve observability across the platform—metrics, logs, traces, alerting, and runbooks
  • Help build custom Kubernetes operators and CRDs to make stateful search clusters declarative, self-healing, and safely upgradable
  • Implement auto-scaling solutions including HPA, VPA, KEDA, and cluster auto-scaling
  • Plan capacity, performance, and cost management across on-premises and cloud infrastructure
  • Participate in on-call rotation to ensure platform reliability

What We're Looking For

Required qualifications:

  • Production Kubernetes experience with hands-on cluster setup and maintenance on your own servers (kubeadm, RKE2, k3s, or similar)—managed-only cloud experience is not sufficient
  • Demonstrated SRE practices including SLOs, error budgets, incident management, and on-call responsibilities
  • Hands-on experience with GitOps or comparable infrastructure/deployment automation; Argo CD or Flux experience is a strong plus
  • Solid observability skills across metrics, logs, traces, and actionable alerting
  • Strong automation mindset—you prioritize fixing root causes over repetitive manual work
  • Collaborative, enabling approach to SRE work—you view infrastructure as a service to developers and balance trade-offs transparently
  • Fluent English required

Valuable but not required:

  • Experience with Harvester, KubeVirt, vSphere/ESXi, OpenStack, or similar virtualization platforms
  • Container storage systems (Longhorn, Ceph) and datacenter networking (load balancing, ingress, VLAN)
  • Auto-scaling implementations and capacity/cost planning
  • Building Kubernetes operators or custom resource definitions
  • German language skills
  • Certifications such as CKA or CKS (welcome but not a substitute for hands-on production experience)

If you have owned production systems, handled incidents, and worked deeply with Kubernetes but your title was never "SRE," we still encourage you to apply. Production experience and an engineering mindset matter more than job titles.

How We Work With You

Candidates apply directly through Get A Job.ai. Our recruiting team will conduct an initial screening to understand your background and match you with the right opportunity. Once we confirm mutual fit, we submit your profile to our client and coordinate the interview process on your behalf.

The client's process includes an introductory call, a take-home technical task (approximately two hours), a 90-minute technical interview with developers, a leadership conversation, and an opportunity to meet the team. Please do not contact the client directly—all communication flows through Get A Job.ai to ensure a smooth, professional experience.

Work model: Hybrid in Berlin (three office days per week)

Tech stack: Kubernetes on own servers, Harvester (KubeVirt), Argo CD/Flux, Prometheus/Grafana, Longhorn/Ceph

Pay

Compensation details will be discussed during the screening process based on your experience level and the client's budget. Our client offers competitive compensation for the European market alongside opportunities for meaningful technical impact and professional growth.

Equal Employment Opportunity: Get A Job.ai is committed to providing equal employment opportunities to all applicants regardless of race, color, religion, sex, national origin, age, disability, or any other protected characteristic. We evaluate candidates based on qualifications and professional merit.

Apply with Get A Job.ai

A recruiter will review your profile and submit you to the client. Do not contact the client directly.

More options

Apply with Get A Job.ai

Apply through Get A Job.ai. A recruiter will review your profile and submit you. Do not contact the client directly.

Apply through Get A Job.ai. A recruiter will review your profile and submit you.

Local insights for this role are preparing — this section updates automatically in a few seconds (or refresh).

Listing facts

  • Role Senior Site Reliability Engineer / SRE – Kubernetes & Hybrid Cloud (m/f/d)
  • Employer Get A Job.ai
  • Location Berlin
  • Type Full Time
  • Pay (from listing) Pay not listed
  • Posted September 8, 2026
  • Apply by October 8, 2026
  • Country Germany
  • Overview Full job description on this page (627 words)

Facts above come from this job record on Get A Job.AI — not copied from third-party review sites.

Limited public data for this employer

We only show facts we can ground in public sources (Wikidata, O*NET, news/discussion links, or this listing). We do not invent Glassdoor-style ratings, salaries, or testimonials when data is thin. Use the listing facts, occupation context, and related openings below while we continue researching.

Explore related openings

Keep exploring on Get A Job.ai

Not quite the right fit? Your next opportunity is a click away.

Hiring instead? Post a job and reach candidates searching right now.

Senior Site Reliability Engineer / SRE – Kuberne… Get A Job.ai · Berlin