- Company: Get A Job.ai
- Location: Berlin
- Salary: Pay not listed
Website Get A Job.ai
Represented by Get A Job.ai
About This Opportunity
We are representing a confidential e-commerce technology organization seeking a Senior Site Reliability Engineer to join their Berlin-based team. This is a rare opportunity to build a modern private cloud platform from the ground up—on your own hardware, not just managed consoles. Our client runs Kubernetes on bare metal in Frankfurt, layering Harvester, Argo CD, and Prometheus to create an on-premises foundation with elastic cloud burst capability. You'll help define SRE practice alongside development teams, with a Team Lead SRE hire coming next to grow the function. The platform you'll operate powers product discovery for thousands of European retail sites processing billions of queries annually—downtime means immediate revenue loss for merchants.
Responsibilities
- Define and own service-level objectives, indicators, and error budgets to drive data-informed reliability decisions across two product stacks
- Lead end-to-end incident response: rapid detection, clear communication during outages, blameless postmortems, and structural fixes that prevent entire classes of future incidents
- Eliminate operational toil through automation and GitOps workflows; evolve observability across metrics, logs, traces, alerting, and runbooks
- Contribute to custom Kubernetes operator development (CRDs) that makes stateful search clusters declarative, self-healing, and safely upgradable
- Implement and tune auto-scaling mechanisms (HPA, VPA, KEDA, cluster auto-scaler) to handle variable catalog sizes and seasonal traffic peaks
- Plan capacity, performance, and cost across on-premises infrastructure and cloud burst scenarios, using AI-assisted tooling where it measurably improves diagnosis speed
- Join on-call rotation with structured buddy support; ship visible reliability improvements within your first 90 days
What We're Looking For
Must-Have Skills:
- Production Kubernetes experience building and maintaining clusters on your own servers (kubeadm, RKE2, k3s or similar)—managed-only experience is not sufficient for this role
- Demonstrated SRE practice: SLOs, error budgets, incident management, and on-call experience
- Hands-on GitOps or comparable infrastructure/deployment automation; Argo CD or Flux experience is a strong advantage
- Solid observability skills with metrics, logs, traces, and reliable alerting systems
- Strong automation instinct—you prefer fixing root causes over repeating workarounds
- Collaborative, enabling mindset: you view SRE as a service to developers, openly discuss trade-offs, and adapt solutions to actual needs
- Fluent English required
Valued But Optional:
- Experience with Harvester, KubeVirt, vSphere/ESXi, OpenStack, or similar virtualization platforms
- Container storage systems (Longhorn, Ceph) and datacenter networking (load balancing, ingress, VLAN)
- Auto-scaling implementation (HPA, VPA, KEDA, cluster auto-scaler) and capacity/cost planning
- Kubernetes operator or CRD development experience
- German language skills
- CKA or CKS certifications (welcome but not a substitute for hands-on production work)
If you've owned production systems, managed incidents, and worked deeply with Kubernetes, apply even if your title was never "SRE"—production experience and engineering mindset matter more than labels.
How We Work With You
Candidates apply through Get A Job.ai. Our talent team conducts an initial screening, then coordinates your interview process with the client: an introduction call, a take-home task (approximately two hours), a 90-minute technical interview with developers, a leadership conversation, and a team meet. We submit qualified candidates directly to the client. Please do not contact the employer independently—all communication flows through our recruiting team to ensure a structured, professional experience.
Location & Work Model
Berlin, hybrid with three office days per week. The role reports to the CTPO initially, transitioning to the incoming Team Lead SRE as the team grows.
Equal Opportunity
Get A Job.ai is committed to inclusive recruiting. We welcome applications from all qualified candidates regardless of background, identity, or circumstance.
Apply with Get A Job.ai
A recruiter will review your profile and submit you to the client. Do not contact the client directly.
Apply through Get A Job.ai. A recruiter will review your profile and submit you. Do not contact the client directly.
Apply through Get A Job.ai. A recruiter will review your profile and submit you.
Local insights for this role are preparing — this section updates automatically in a few seconds (or refresh).
Listing facts
- Role Senior Site Reliability Engineer / SRE – Kubernetes & Hybrid Cloud (m/f/d)
- Employer Get A Job.ai
- Location Berlin
- Type Full Time
- Pay (from listing) Pay not listed
- Posted September 17, 2026
- Apply by October 17, 2026
- Country Germany
- Overview Full job description on this page (568 words)
Facts above come from this job record on Get A Job.AI — not copied from third-party review sites.
Limited public data for this employer
We only show facts we can ground in public sources (Wikidata, O*NET, news/discussion links, or this listing). We do not invent Glassdoor-style ratings, salaries, or testimonials when data is thin. Use the listing facts, occupation context, and related openings below while we continue researching.
Explore related openings
Keep exploring on Get A Job.ai
Not quite the right fit? Your next opportunity is a click away.
- Browse all jobs
- More jobs by category
- Remote jobs you can do from anywhere
- Research typical pay for this role
- Set a job alert so new matches reach you first
- Upload your resume to apply faster
Hiring instead? Post a job and reach candidates searching right now.
