Loading...

Site Reliability Engineer (m/w/d)

  • Company: Get A Job.ai
  • Location: Köln, Nordrhein-Westfalen, Deutschland
  • Salary: Pay not listed

Website Get A Job.ai

Represented by Get A Job.ai

Responsibilities

We are representing a confidential cloud infrastructure organization based in Cologne, and our talent team is seeking a Site Reliability Engineer to join their growing platform engineering initiative.

In this role, you will help build and industrialize an on-premise cloud platform based on OpenStack. Working within a small, experienced team, you'll contribute to both the core infrastructure and the Kubernetes/GitOps stack that powers their customer-facing platform. AI-assisted engineering is an integral part of their daily practice—from spec-driven development to incident response and automation.

Your responsibilities will include:

  • Designing and developing OpenStack-based on-premise cloud infrastructure with the goal of highly automated bare-metal deployment and operations
  • Building and operating Infrastructure as Code using Ansible and Terraform, plus Kubernetes and GitOps workflows with FluxCD/ArgoCD—supported by LLMs, agentic workflows, and automated testing
  • Managing the full lifecycle of compute infrastructure—from bare metal provisioning, firmware, and hypervisors to patching, migrations, host evacuations, capacity rebalancing, and operational automation
  • Advancing AI substrate and self-healing strategies, including structured knowledge bases, agentic workflows for incident triage and capacity planning, and progressive automation of runbooks
  • Designing tests for non-regression, performance, and security; documenting and packaging solutions; continuously improving the platform based on telemetry, operational experience, and user feedback
  • Serving as a technical point of contact and sparring partner for colleagues on automation, platform engineering, and AI tooling

What We're Looking For

Our client seeks a senior-level professional with several years of hands-on experience as an SRE, Platform Engineer, or DevOps Engineer operating production infrastructure. You should bring:

  • Deep practical experience with OpenStack, Kubernetes, and Linux, including bare-metal environments
  • End-to-end compute infrastructure experience—firmware/BIOS rollouts, bare-metal provisioning, hardware diagnostics, hypervisors, migrations, host evacuations, graceful drains, and capacity rebalancing
  • Active use of AI-assisted engineering in your daily work. You deploy LLMs and agentic tools strategically where they genuinely support development, testing, reviews, or operations, and you can assess where AI adds real value versus where solid engineering expertise remains critical
  • Proficiency with Ansible, Terraform, and GitOps workflows (FluxCD or ArgoCD), with experience running automated processes reliably in production
  • Experience with Go and/or Python, plus agentic coding environments such as Claude Code, Cursor, Aider, or similar
  • Familiarity with observability, networking, compute tuning, auto-remediation, security-critical infrastructure, and multi-site cloud environments
  • A strong ownership mindset and the drive to not just operate systems but continuously improve them. You enjoy sharing knowledge and can clearly communicate complex technical topics
  • Comfort working in an international environment in English, including technical discussions, documentation, and team collaboration

Technical stack: OpenStack, Kubernetes, KVM, Linux, Bare Metal, Ansible, Terraform, Go, FluxCD/ArgoCD, Git, Python, Claude Code, Cursor, and agentic coding tooling.

How We Work with You

Candidates apply directly through Get A Job.ai. Our recruiting team will screen your application and conduct an initial conversation to understand your background and interest. If there's a strong match, we will submit your profile to our client for consideration. Please do not contact the client directly—all communication and coordination will flow through Get A Job.ai to ensure a smooth and professional process.

Pay

Pay details will be discussed during the screening process and are competitive with the market for senior platform engineering roles in Germany.

Equal Employment Opportunity: Get A Job.ai is committed to inclusive hiring practices. We welcome applications from all qualified candidates regardless of race, color, religion, sex, national origin, age, disability, or any other protected characteristic.

Apply with Get A Job.ai

A recruiter will review your profile and submit you to the client. Do not contact the client directly.

More options

Apply with Get A Job.ai

Apply through Get A Job.ai. A recruiter will review your profile and submit you. Do not contact the client directly.

Apply through Get A Job.ai. A recruiter will review your profile and submit you.

Local insights for this role are preparing — this section updates automatically in a few seconds (or refresh).

Listing facts

  • Role Site Reliability Engineer (m/w/d)
  • Employer Get A Job.ai
  • Location Köln, Nordrhein-Westfalen, Deutschland
  • Type Full Time
  • Pay (from listing) Pay not listed
  • Posted September 11, 2026
  • Apply by October 11, 2026
  • Country Germany
  • Overview Full job description on this page (566 words)

Facts above come from this job record on Get A Job.AI — not copied from third-party review sites.

Typical work in Senior Site Reliability Engineer

Independent occupational context from O*NET (U.S. public-domain labor data). This is about the occupation, not a rewrite of this employer's posting.

  • Study product characteristics or customer requirements to determine validation objectives and standards.
  • Analyze validation test data to determine whether systems or processes have met validation criteria or to identify root causes of production problems.
  • Develop validation master plans, process flow diagrams, test cases, or standard operating procedures.
  • Prepare detailed reports or design statements, based on results of validation and qualification tests or reviews of procedures and protocols.
  • Maintain validation test equipment.
  • Conduct validation or qualification tests of new or existing processes, equipment, or software in accordance with internal protocols or external standards.

Source: O*NET

Limited public data for this employer

We only show facts we can ground in public sources (Wikidata, O*NET, news/discussion links, or this listing). We do not invent Glassdoor-style ratings, salaries, or testimonials when data is thin. Use the listing facts, occupation context, and related openings below while we continue researching.

Explore related openings

Occupation family: Senior Site Reliability Engineer

Keep exploring on Get A Job.ai

Not quite the right fit? Your next opportunity is a click away.

Hiring instead? Post a job and reach candidates searching right now.

Site Reliability Engineer (m/w/d) Get A Job.ai · Köln, Nordrhein-Westfalen, Deutschland