Loading...

Senior Site Reliability Engineer / SRE – Kubernetes & Hybrid Cloud (m/f/d)

  • Company: Get A Job.ai
  • Location: Berlin
  • Salary: Pay not listed
  • Full Time
  • Berlin

Website Get A Job.ai

Represented by Get A Job.ai

About This Opportunity

We are representing a confidential eCommerce technology organization seeking a Senior Site Reliability Engineer to join their growing team in Berlin. This is a hybrid role working directly with development teams to build and evolve a modern private cloud platform.

What makes this position unique: you'll work with actual on-premises infrastructure running Kubernetes and Harvester, not just managed cloud services. Our client operates their own hardware in Frankfurt and is building toward a hybrid cloud architecture that balances on-prem reliability with elastic cloud burst capability. You'll help define SRE practices from the ground up, with a Team Lead SRE hire planned to join soon.

The platform you'll support powers product discovery for over 2,000 European online retailers, handling billions of customer queries annually. Reliability directly impacts client revenue.

Responsibilities

  • Define and maintain SLOs, SLIs and error budgets to drive data-informed reliability decisions
  • Lead incident response end-to-end: detection, communication, blameless postmortems, and structural fixes that prevent entire classes of future incidents
  • Eliminate operational toil through automation and GitOps practices
  • Evolve observability across metrics, logs, traces, alerting and runbooks for multiple technology stacks
  • Contribute to building custom Kubernetes operators (CRDs) that make stateful search clusters declarative, self-healing and safely upgradable
  • Implement and tune auto-scaling (HPA/VPA, KEDA, cluster autoscaler) for infrastructure that currently lacks it
  • Plan capacity, performance and cost across on-premises and cloud environments, including peak-season load scenarios
  • Join on-call rotation with team support during your onboarding period

What We're Looking For

Must-have qualifications:

  • Production Kubernetes experience building clusters on bare metal or your own servers (kubeadm, RKE2, k3s or similar) – you understand cluster lifecycle and upgrades, not just consumption of managed services
  • Lived SRE practice: hands-on experience with SLOs, error budgets, incident management and on-call responsibilities
  • GitOps or comparable infrastructure/deployment automation experience; Argo CD or Flux experience is a strong advantage
  • Solid observability skills across metrics, logs, traces and reliable alerting
  • Strong automation instinct – you prefer fixing root causes over repeated manual interventions
  • Collaborative, enabling mindset – you view SRE as a service to developers, asking what they need and discussing trade-offs openly
  • Fluent English required for team collaboration

Nice-to-have experience (genuinely optional):

  • Harvester, KubeVirt, vSphere/ESXi, OpenStack or similar virtualization/hyperconverged platforms
  • Container storage solutions (Longhorn, Ceph) and datacenter networking (load balancing, ingress, VLAN)
  • Auto-scaling technologies (HPA, VPA, KEDA, cluster autoscaler) and capacity/cost planning
  • Building Kubernetes operators or custom resource definitions
  • German language skills
  • Certifications (CKA, CKS) – welcome but not a substitute for hands-on production experience

If your title was never "SRE" but you've owned production systems, handled incidents and worked deeply with Kubernetes, we encourage you to apply. Production experience and engineering mindset matter more than job titles.

How We Work With You

Candidates apply directly through Get A Job.ai. Our recruiting team will conduct an initial screening to understand your background and match it to our client's requirements. Strong candidates are then submitted to the client for their interview process, which includes a technical take-home task (approximately 2 hours), a 90-minute technical interview with developers, a leadership conversation, and a team meet.

Please do not contact the client directly. All communication and coordination goes through our talent team to ensure a smooth, professional process for everyone involved.

Position Details

Location: Berlin, hybrid work model with three office days per week

Tech stack: Kubernetes on own servers, Harvester (KubeVirt), Argo CD/Flux, Prometheus/Grafana, Longhorn/Ceph

Pay: Compensation details will be discussed during the screening process based on your experience level.

Get A Job.ai is committed to equal employment opportunity regardless of race, color, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, or veteran status.

Apply with Get A Job.ai

A recruiter will review your profile and submit you to the client. Do not contact the client directly.

More options

Apply with Get A Job.ai

Apply through Get A Job.ai. A recruiter will review your profile and submit you. Do not contact the client directly.

Apply through Get A Job.ai. A recruiter will review your profile and submit you.

Terms used in this posting

on-call
You may be required to be reachable and available to work outside normal scheduled hours, typically for a set rotation.
hybrid
A work arrangement combining both in-office and remote/at-home work, typically on a set schedule.

Working in Berlin, Deutschland

Weather right now in Berlin, Deutschland: checking… · Local time: · Air quality: · Daylight: · UV index: · Wind: · Pollen:

Berlin is the capital of Germany as well as its largest city by both area and population. With 3.7 million inhabitants, it has the highest population within its city limits of any city in the European Union. The city is also one of the states of Germany, being the third-smallest state in the country by area. Berlin is surrounded by the state of Brandenburg, bordering Brandenburg's capital Potsdam to the southwest. The urban area of Berlin has a population of over 5 million, making it the most populous in Germany. The Berlin-Brandenburg capital region has around 6 million inhabitants and is Ger

Note: Germany observes a public holiday on Oct 3 — German Unity Day.

🇩🇪 Relocation safety for Germany: Very Safevia Warnely, CC BY 4.0

National unemployment rate in Germany: 3.7%via World Bank

GDP per capita in Germany: $60,496via World Bank

Consumer price inflation in Germany: 2.2% (annual) — via World Bank

Real GDP growth in Germany: 0.2% (annual) — via World Bank

Statutory minimum wage in Germany: €2,343/monthvia Eurostat

Cost of living in Germany: 9.1% above the EU averagevia Eurostat

Job vacancy rate in Germany: 2.8%via Eurostat

Average hours worked per year in Germany: 1,332via OECD

Nearby green space: 10 parks within 1.5km — closest is Lustgarten (306m). via OpenStreetMap

Nearest public transit: Staatsoper (bus stop, 55m). via OpenStreetMap

  • Elevation 36m (118 ft)

Job details above are provided by the employer/source. The sections on this page are compiled from public data sources with AI assistance.

Accommodations: if you need a workplace accommodation to apply for or perform this job, see ADA.gov or EEOC.gov for guidance on your rights and how to request one.

Add application deadline to calendar

Listing facts

  • Role Senior Site Reliability Engineer / SRE – Kubernetes & Hybrid Cloud (m/f/d)
  • Employer Get A Job.ai
  • Location Berlin
  • Type Full Time
  • Pay (from listing) Pay not listed
  • Posted September 12, 2026
  • Apply by October 12, 2026
  • Country Germany
  • Overview Full job description on this page (605 words)

Facts above come from this job record on Get A Job.AI — not copied from third-party review sites.

Limited public data for this employer

We only show facts we can ground in public sources (Wikidata, O*NET, news/discussion links, or this listing). We do not invent Glassdoor-style ratings, salaries, or testimonials when data is thin. Use the listing facts, occupation context, and related openings below while we continue researching.

Explore related openings

Keep exploring on Get A Job.ai

Not quite the right fit? Your next opportunity is a click away.

Hiring instead? Post a job and reach candidates searching right now.

Senior Site Reliability Engineer / SRE – Kuberne… Get A Job.ai · Berlin