- Company: Get A Job.ai
- Location: Leipzig
- Salary: Pay not listed
- Work type: Remote
Website Get A Job.ai
Represented by Get A Job.ai
About This Opportunity
We are representing a well-funded AI technology company based in Leipzig that is building enterprise-grade solutions for organizations handling complex data and mission-critical workflows. As their customer base and engineering capabilities expand, they are making a significant investment in platform infrastructure and operations.
Our client is seeking their first dedicated Senior DevOps / Platform Engineer to take technical ownership of production infrastructure and shape the next evolution of their platform. This is a hands-on senior individual contributor role where you will work directly with software engineers and company leadership to operate and evolve a hybrid infrastructure environment spanning self-administered Linux servers, European infrastructure providers, dedicated hardware, and multi-cloud services.
This position is fully remote within Germany, with regular opportunities to collaborate with the team in Leipzig (travel covered). Given the autonomy and technical responsibility involved, our client requires candidates with a proven track record of independently operating business-critical production systems.
Responsibilities
You will take technical ownership of production infrastructure, including:
- Operating and evolving self-administered Linux systems across virtual machines and dedicated servers, including containers, networks, reverse proxies, and API gateways
- Making deployments safe, repeatable, and developer-friendly through Infrastructure as Code, CI/CD workflows, automated checks, versioning, and rollback capabilities
- Operating and improving PostgreSQL in production, including performance analysis, connection pooling, capacity planning, backups, and regularly tested restore procedures
- Developing the self-hosted observability stack to connect metrics, logs, traces, and actionable alerts
- Strengthening security across infrastructure through IAM, least privilege, secrets management, TLS, vulnerability scanning, and patch management
- Implementing technical controls for ISO 27001 with continuous, auditable evidence generation
- Shaping AI infrastructure by integrating and evaluating model and inference providers based on reliability, latency, throughput, cost, and operational effort
- Exploring self-hosted LLM inference opportunities, potentially including GPU infrastructure and serving technologies
- Improving incident response and operational resilience through root cause analysis, runbooks, and documentation
- Enhancing internal developer experience by reducing manual work and creating clear workflows
- Making infrastructure and inference costs transparent to inform build, buy, and hosting decisions
You will help determine initial priorities after assessing the existing platform, identifying relevant risks, explaining available options, and taking improvements through to reliable production operation.
What We're Looking For
Required Technical Experience:
- Several years operating production SaaS systems on Linux servers you or your team administered directly (experience limited to fully managed cloud services is not sufficient)
- Production experience with Docker and Docker Compose, plus reverse proxies or API gateways such as Traefik or Kong
- Practical PostgreSQL operations experience, including backups and restores you have personally configured and tested, performance analysis, and connection pooling
- Experience building and maintaining CI/CD pipelines using GitHub Actions, GitLab CI, or comparable systems
- Infrastructure as Code experience with Terraform, Pulumi, or comparable tooling
- Experience operating observability stacks using tools such as Grafana, Loki, Prometheus, Sentry, or comparable technologies
- Understanding of production infrastructure security foundations, including IAM, least privilege, secrets management, vulnerability scanning, and patching
- Ability to design infrastructure that is reliable for customers and straightforward for engineers to use
- Fluent English communication (German helpful but not required)
AI Infrastructure Foundation:
- Understanding of how RAG systems work and where they commonly fail in production
- Ability to discuss trade-offs in LLM serving, including latency, throughput, reliability, cost, and operational complexity
- Following developments in models, inference providers, and the GenAI market closely enough to evaluate new options critically
- Strong platform engineering foundation with technical curiosity to develop deeper AI infrastructure expertise (prior GPU operations or self-hosted inference experience helpful but not required)
How You Work:
- Take responsibility from initial investigation through production operation
- Investigate unfamiliar systems until you understand relevant behavior and failure modes
- Make architecture decisions deliberately and document the reasoning behind them
- Value maintainability, automation, tested recovery, and clear operational processes
- Prioritize according to risk and impact and communicate trade-offs openly
- Remain structured during incidents and communicate clearly with technical and non-technical stakeholders
- Use AI tools productively while remaining accountable for every architecture decision and production change
- Communicate proactively in a distributed team and make progress, decisions, risks, and blockers visible
For This Remote Role:
Given the broad technical responsibility and high degree of autonomy, our client requires evidence that you have already:
- Independently operated business-critical production systems
- Identified and prioritized infrastructure work without waiting for detailed instructions
- Led technical improvements across team boundaries
- Communicated reliably in writing and remotely
- Remained accountable after deployment of changes
Helpful But Not Required:
- Experience with Hetzner, NixOS, or self-hosted Supabase
- Experience implementing technical ISO 27001 controls and audit evidence
- Experience with SRE practices such as SLOs, error budgets, and incident reviews
- Experience operating GPUs or serving LLMs with technologies such as vLLM
Important: This role is open exclusively to candidates currently based in Germany.
How We Work With You
When you apply through Get A Job.ai, our talent team will review your application and conduct an initial screening. We will then present qualified candidates to our client for consideration. The client's hiring process consists of three stages: a 30-minute introductory conversation, a 90-minute technical interview, and a final conversation with company leadership. They move quickly and keep candidates informed throughout.
Please do not attempt to contact the client directly. All applications must go through Get A Job.ai to be considered.
Along with your CV, please provide brief answers to these three questions:
- What is the most critical production system you have personally operated, and what were you directly responsible for?
- Tell us about one infrastructure risk, incident, or reliability problem you identified and improved.
- Why does this remote role and its level of ownership fit the way you work?
Please write these answers yourself—our client is not assessing polished marketing language but wants to understand your actual experience, judgment, and motivation.
What's Offered
- Substantial technical area to shape and own in a well-funded, growing organization
- Direct influence on architecture, reliability, security, and engineering productivity
- Opportunity to deepen expertise in AI infrastructure and LLM workload operations
- Close collaboration with engineering team, CTO, and company leadership
- Short decision paths and authority to move important infrastructure work forward
- Permanent, full-time position with flexible working hours
- Fully remote work from anywhere within Germany
- Regular opportunities to meet and work with the team in Leipzig, with travel and accommodation covered
- Access to WellPass fitness membership
- Opportunity to participate in VSOP
Pay
The client offers €75,000–€85,000 gross annual salary, depending on experience and scope.
Get A Job.ai is an equal opportunity recruiter. We welcome applications from all qualified candidates regardless of race, color, religion, sex, national origin, age, disability, or any other protected characteristic.
Apply with Get A Job.ai
A recruiter will review your profile and submit you to the client. Do not contact the client directly.
Apply through Get A Job.ai. A recruiter will review your profile and submit you. Do not contact the client directly.
Apply through Get A Job.ai. A recruiter will review your profile and submit you.
Terms used in this posting
- hybrid
- A work arrangement combining both in-office and remote/at-home work, typically on a set schedule.
Explore Get A Job.ai online
Working in Leipzig, Leipzig (Kreis)
Weather right now in Leipzig, Leipzig (Kreis): checking… · Local time: · Air quality: · Daylight: · UV index: · Wind: · Pollen:
Note: Germany observes a public holiday on Oct 3 — German Unity Day.
🇩🇪 Relocation safety for Germany: Very Safe — via Warnely, CC BY 4.0
National unemployment rate in Germany: 3.7% — via World Bank
GDP per capita in Germany: $60,496 — via World Bank
Consumer price inflation in Germany: 2.2% (annual) — via World Bank
Real GDP growth in Germany: 0.2% (annual) — via World Bank
Statutory minimum wage in Germany: €2,343/month — via Eurostat
Cost of living in Germany: 9.1% above the EU average — via Eurostat
Job vacancy rate in Germany: 2.8% — via Eurostat
Average hours worked per year in Germany: 1,332 — via OECD
Nearby green space: 9 parks within 1.5km — closest is Thomaswiese (123m). via OpenStreetMap
Nearest public transit: Markt (transit stop, 89m). via OpenStreetMap
- Elevation 117m (384 ft)
Job details above are provided by the employer/source. The sections on this page are compiled from public data sources with AI assistance.
Accommodations: if you need a workplace accommodation to apply for or perform this job, see ADA.gov or EEOC.gov for guidance on your rights and how to request one.
Listing facts
- Role Senior DevOps / Platform Engineer, AI Infrastructure (m/f/x)
- Employer Get A Job.ai
- Location Leipzig · Remote-friendly
- Type Full Time
- Pay (from listing) Pay not listed
- Posted September 15, 2026
- Apply by October 15, 2026
- Country Germany
- Overview Full job description on this page (1,081 words)
Facts above come from this job record on Get A Job.AI — not copied from third-party review sites.
Limited public data for this employer
We only show facts we can ground in public sources (Wikidata, O*NET, news/discussion links, or this listing). We do not invent Glassdoor-style ratings, salaries, or testimonials when data is thin. Use the listing facts, occupation context, and related openings below while we continue researching.
Explore related openings
Keep exploring on Get A Job.ai
Not quite the right fit? Your next opportunity is a click away.
- Browse all jobs
- More jobs by category
- Remote jobs you can do from anywhere
- Research typical pay for this role
- Set a job alert so new matches reach you first
- Upload your resume to apply faster
Hiring instead? Post a job and reach candidates searching right now.
