Loading...

Senior DevOps / Platform Engineer, AI Infrastructure (m/f/x)

  • Company: Get A Job.ai
  • Location: Leipzig
  • Salary: Pay not listed
  • Work type: Remote

Website Get A Job.ai

Represented by Get A Job.ai

About This Opportunity

We are representing a well-funded AI technology company based in Leipzig that is building enterprise-grade solutions for organizations handling complex data and mission-critical workflows. As their customer base and engineering capabilities expand, they are making a significant investment in platform infrastructure and operations.

Our client is seeking their first dedicated Senior DevOps / Platform Engineer to take technical ownership of production infrastructure and shape the next evolution of their platform. This is a hands-on senior individual contributor role where you will work directly with software engineers and company leadership to operate and evolve a hybrid infrastructure environment spanning self-administered Linux servers, European infrastructure providers, dedicated hardware, and multi-cloud services.

This position is fully remote within Germany, with regular opportunities to collaborate with the team in Leipzig (travel covered). Given the autonomy and technical responsibility involved, our client requires candidates with a proven track record of independently operating business-critical production systems.

Responsibilities

You will take technical ownership of production infrastructure, including:

  • Operating and evolving self-administered Linux systems across virtual machines and dedicated servers, including containers, networks, reverse proxies, and API gateways
  • Making deployments safe, repeatable, and developer-friendly through Infrastructure as Code, CI/CD workflows, automated checks, versioning, and rollback capabilities
  • Operating and improving PostgreSQL in production, including performance analysis, connection pooling, capacity planning, backups, and regularly tested restore procedures
  • Developing the self-hosted observability stack to connect metrics, logs, traces, and actionable alerts
  • Strengthening security across infrastructure through IAM, least privilege, secrets management, TLS, vulnerability scanning, and patch management
  • Implementing technical controls for ISO 27001 with continuous, auditable evidence generation
  • Shaping AI infrastructure by integrating and evaluating model and inference providers based on reliability, latency, throughput, cost, and operational effort
  • Exploring self-hosted LLM inference opportunities, potentially including GPU infrastructure and serving technologies
  • Improving incident response and operational resilience through root cause analysis, runbooks, and documentation
  • Enhancing internal developer experience by reducing manual work and creating clear workflows
  • Making infrastructure and inference costs transparent to inform build, buy, and hosting decisions

You will help determine initial priorities after assessing the existing platform, identifying relevant risks, explaining available options, and taking improvements through to reliable production operation.

What We're Looking For

Required Technical Experience:

  • Several years operating production SaaS systems on Linux servers you or your team administered directly (experience limited to fully managed cloud services is not sufficient)
  • Production experience with Docker and Docker Compose, plus reverse proxies or API gateways such as Traefik or Kong
  • Practical PostgreSQL operations experience, including backups and restores you have personally configured and tested, performance analysis, and connection pooling
  • Experience building and maintaining CI/CD pipelines using GitHub Actions, GitLab CI, or comparable systems
  • Infrastructure as Code experience with Terraform, Pulumi, or comparable tooling
  • Experience operating observability stacks using tools such as Grafana, Loki, Prometheus, Sentry, or comparable technologies
  • Understanding of production infrastructure security foundations, including IAM, least privilege, secrets management, vulnerability scanning, and patching
  • Ability to design infrastructure that is reliable for customers and straightforward for engineers to use
  • Fluent English communication (German helpful but not required)

AI Infrastructure Foundation:

  • Understanding of how RAG systems work and where they commonly fail in production
  • Ability to discuss trade-offs in LLM serving, including latency, throughput, reliability, cost, and operational complexity
  • Following developments in models, inference providers, and the GenAI market closely enough to evaluate new options critically
  • Strong platform engineering foundation with technical curiosity to develop deeper AI infrastructure expertise (prior GPU operations or self-hosted inference experience helpful but not required)

How You Work:

  • Take responsibility from initial investigation through production operation
  • Investigate unfamiliar systems until you understand relevant behavior and failure modes
  • Make architecture decisions deliberately and document the reasoning behind them
  • Value maintainability, automation, tested recovery, and clear operational processes
  • Prioritize according to risk and impact and communicate trade-offs openly
  • Remain structured during incidents and communicate clearly with technical and non-technical stakeholders
  • Use AI tools productively while remaining accountable for every architecture decision and production change
  • Communicate proactively in a distributed team and make progress, decisions, risks, and blockers visible

For This Remote Role:

Given the broad technical responsibility and high degree of autonomy, our client requires evidence that you have already:

  • Independently operated business-critical production systems
  • Identified and prioritized infrastructure work without waiting for detailed instructions
  • Led technical improvements across team boundaries
  • Communicated reliably in writing and remotely
  • Remained accountable after deployment of changes

Helpful But Not Required:

  • Experience with Hetzner, NixOS, or self-hosted Supabase
  • Experience implementing technical ISO 27001 controls and audit evidence
  • Experience with SRE practices such as SLOs, error budgets, and incident reviews
  • Experience operating GPUs or serving LLMs with technologies such as vLLM

Important: This role is open exclusively to candidates currently based in Germany.

How We Work With You

When you apply through Get A Job.ai, our talent team will review your application and conduct an initial screening. We will then present qualified candidates to our client for consideration. The client's hiring process consists of three stages: a 30-minute introductory conversation, a 90-minute technical interview, and a final conversation with company leadership. They move quickly and keep candidates informed throughout.

Please do not attempt to contact the client directly. All applications must go through Get A Job.ai to be considered.

Along with your CV, please provide brief answers to these three questions:

  • What is the most critical production system you have personally operated, and what were you directly responsible for?
  • Tell us about one infrastructure risk, incident, or reliability problem you identified and improved.
  • Why does this remote role and its level of ownership fit the way you work?

Please write these answers yourself—our client is not assessing polished marketing language but wants to understand your actual experience, judgment, and motivation.

What's Offered

  • Substantial technical area to shape and own in a well-funded, growing organization
  • Direct influence on architecture, reliability, security, and engineering productivity
  • Opportunity to deepen expertise in AI infrastructure and LLM workload operations
  • Close collaboration with engineering team, CTO, and company leadership
  • Short decision paths and authority to move important infrastructure work forward
  • Permanent, full-time position with flexible working hours
  • Fully remote work from anywhere within Germany
  • Regular opportunities to meet and work with the team in Leipzig, with travel and accommodation covered
  • Access to WellPass fitness membership
  • Opportunity to participate in VSOP

Pay

The client offers €75,000–€85,000 gross annual salary, depending on experience and scope.

Get A Job.ai is an equal opportunity recruiter. We welcome applications from all qualified candidates regardless of race, color, religion, sex, national origin, age, disability, or any other protected characteristic.

Apply with Get A Job.ai

A recruiter will review your profile and submit you to the client. Do not contact the client directly.

More options

Apply with Get A Job.ai

Apply through Get A Job.ai. A recruiter will review your profile and submit you. Do not contact the client directly.

Apply through Get A Job.ai. A recruiter will review your profile and submit you.

Local insights for this role are preparing — this section updates automatically in a few seconds (or refresh).

Listing facts

  • Role Senior DevOps / Platform Engineer, AI Infrastructure (m/f/x)
  • Employer Get A Job.ai
  • Location Leipzig · Remote-friendly
  • Type Full Time
  • Pay (from listing) Pay not listed
  • Posted September 15, 2026
  • Apply by October 15, 2026
  • Country Germany
  • Overview Full job description on this page (1,081 words)

Facts above come from this job record on Get A Job.AI — not copied from third-party review sites.

Limited public data for this employer

We only show facts we can ground in public sources (Wikidata, O*NET, news/discussion links, or this listing). We do not invent Glassdoor-style ratings, salaries, or testimonials when data is thin. Use the listing facts, occupation context, and related openings below while we continue researching.

Explore related openings

Keep exploring on Get A Job.ai

Not quite the right fit? Your next opportunity is a click away.

Hiring instead? Post a job and reach candidates searching right now.

Senior DevOps / Platform Engineer, AI Infrastruc… Get A Job.ai · Leipzig