- Company: Get A Job.ai
- Location: Germany
- Salary: Pay not listed
- Work type: Remote
Website Get A Job.ai
Represented by Get A Job.ai
About This Opportunity
We are representing a confidential cloud infrastructure organization building a European AI Infrastructure & Platform Operations team. This team operates large-scale AI environments powered by NVIDIA GPUs, high-performance networking, Kubernetes, and next-generation platform technologies across multiple datacenters.
Working at the intersection of infrastructure, networking, and platform operations, you will help ensure the availability, performance, and operational stability of critical AI infrastructure platforms that power modern AI workloads. This is an opportunity to work with cutting-edge technologies in AI infrastructure while contributing to the evolution of AI-powered operational services.
This is a remote position available to candidates based in Germany or elsewhere in the EU.
Responsibilities
- Monitor, operate, and support production AI infrastructure platforms
- Investigate and resolve infrastructure, networking, hardware, and platform-related incidents
- Support NVIDIA GPU infrastructure and associated platform services
- Monitor and troubleshoot Kubernetes-based environments
- Investigate performance, availability, and reliability issues across infrastructure and platform components
- Collaborate with engineering teams, hardware vendors, datacenter personnel, and service delivery teams to resolve technical issues
- Participate in incident response, root cause analysis, and operational improvement activities
- Contribute to improvements in monitoring, observability, automation, and operational processes
- Maintain operational documentation, runbooks, and knowledge articles
What We're Looking For
Required qualifications:
- 3+ years of experience in infrastructure operations, platform operations, network operations, site reliability engineering, cloud operations, datacenter operations, or related technical roles
- Strong Linux administration and troubleshooting skills
- Good understanding of networking concepts and experience diagnosing infrastructure-related issues
- Working knowledge of Kubernetes in production environments
- Experience supporting production infrastructure and services
- Strong analytical and problem-solving skills
- Experience working within structured operational and incident management processes
- Excellent communication and collaboration skills
- Ability to work within a shift-based operational environment
Highly desirable experience in one or more of the following areas:
- NVIDIA GPU infrastructure and accelerated computing platforms
- InfiniBand networking and NVIDIA UFM
- Kubernetes platform operations
- AI infrastructure or HPC environments
- Site Reliability Engineering (SRE) or Platform Engineering
- Observability platforms such as Grafana, Prometheus, ELK, or OpenTelemetry
- Infrastructure automation technologies and Infrastructure-as-Code practices
- Large-scale distributed systems and production platforms
How We Work With You
Our talent team at Get A Job.ai partners exclusively with this organization to identify qualified candidates. When you apply through our platform, a recruiter will review your profile and qualifications. Strong candidates will be screened by our team before being submitted directly to our client for consideration.
Please apply only through Get A Job.ai. Do not contact the client organization directly, as we manage the entire candidate submission process.
Pay
Compensation details will be discussed during the screening process with our recruiting team.
Equal Opportunity: Get A Job.ai is committed to inclusive hiring practices and works with clients who value diversity and equal opportunity in employment.
Apply with Get A Job.ai
A recruiter will review your profile and submit you to the client. Do not contact the client directly.
Apply through Get A Job.ai. A recruiter will review your profile and submit you. Do not contact the client directly.
Apply through Get A Job.ai. A recruiter will review your profile and submit you.
Local insights for this role are preparing — this section updates automatically in a few seconds (or refresh).
Listing facts
- Role AI Infrastructure & Platform Operations Engineer (remote in the EU)
- Employer Get A Job.ai
- Location Germany
- Type Full Time
- Pay (from listing) Pay not listed
- Posted September 12, 2026
- Apply by October 13, 2026
- Country Germany
- Overview Full job description on this page (449 words)
Facts above come from this job record on Get A Job.AI — not copied from third-party review sites.
Limited public data for this employer
We only show facts we can ground in public sources (Wikidata, O*NET, news/discussion links, or this listing). We do not invent Glassdoor-style ratings, salaries, or testimonials when data is thin. Use the listing facts, occupation context, and related openings below while we continue researching.
Explore related openings
Keep exploring on Get A Job.ai
Not quite the right fit? Your next opportunity is a click away.
- Browse all jobs
- More jobs by category
- Remote jobs you can do from anywhere
- Research typical pay for this role
- Set a job alert so new matches reach you first
- Upload your resume to apply faster
Hiring instead? Post a job and reach candidates searching right now.
