Loading...

Cloud Service Security Platform DevOps & Maintenance

  • Company: Get A Job.ai
  • Location: United States
  • Salary: Pay not listed

Website Get A Job.ai

Represented by Get A Job.ai

About This Opportunity

We are representing a confidential technology client in the United States who is building cutting-edge AI infrastructure at scale. They need a talented DevOps engineer to own the platform substrate that enables model teams, data scientists, and product engineers to ship AI products without friction.

This is an entry-to-mid level role where you'll build the paved road for every AI product the organization ships—the CI/CD pipelines, Internal Developer Platform, and MLOps substrate that lets teams deploy without opening a ticket.

Responsibilities

  • CI/CD & MLOps Pipeline Management: Design, implement, and maintain end-to-end CI/CD pipelines for both software applications and machine learning models. Automate build, test, deployment, and rollback processes to ensure seamless transitions from innovation to production.
  • Cloud-Native & AI Infrastructure: Build, optimize, and scale cloud-native infrastructure using Kubernetes and Docker. Manage and provision specialized computing resources such as GPU clusters to support high-performance AI workloads and model inferencing.
  • High Availability Architecture: Take ownership of high-availability design in production environments. Implement disaster recovery strategies, self-healing mechanisms, capacity planning, and performance tuning to meet stringent business SLAs.
  • Infrastructure as Code: Champion IaC practices utilizing tools such as Terraform, Ansible, and Helm to achieve fully automated, reproducible, and auditable infrastructure provisioning across multiple cloud environments.
  • Observability & Monitoring: Architect and refine comprehensive monitoring, logging, and alerting systems using tools like Prometheus, Grafana, and the ELK/EFK stack to provide deep visibility into system health, application performance, and AI model metrics.
  • Internal Developer Platform: Build the paved road that lets product, model, and data-science teams deploy without opening a ticket. Make golden paths so obvious that shortcuts feel harder than doing it right.
  • Cross-functional Collaboration: Work closely with R&D, Data Science, Security, and Business teams to streamline workflows, eliminate bottlenecks, and continuously elevate engineering efficiency.
  • Governance, Security & Compliance: Establish and enforce system stability and security standards. Manage release workflows, implement Zero Trust access controls, oversee secrets management, and ensure compliance frameworks.
  • Incident Management & Resolution: Act as the technical lead during complex system anomalies and major incidents. Spearhead rapid troubleshooting, thorough root cause analysis, and preventative remediation—turning each incident into an automation that stops the next one before it pages a human.

What We're Looking For

Required Qualifications:

  • Bachelor's degree or above in Computer Science, Engineering, or a related technical field
  • 5+ years of hands-on experience in DevOps, Site Reliability Engineering, or Cloud Infrastructure roles
  • Expert-level knowledge of Linux operating systems and core networking principles (TCP/IP, DNS, HTTP, Load Balancing, VPCs)
  • Deep mastery of Docker and Kubernetes orchestration, including cluster management and production-level best practices
  • Proven proficiency in designing and managing infrastructure on major public or hybrid cloud platforms such as AWS, GCP, or Azure, including multi-cloud and hybrid-cloud strategies
  • Strong coding and scripting capabilities in at least one major language (Go, Python, Shell, etc.) with a solid engineering-oriented mindset focused on automation and tooling development
  • Systematic and practical understanding of CI/CD methodologies, Infrastructure as Code, Observability paradigms, and Site Reliability Engineering principles
  • Exceptional problem-solving abilities, sharp technical judgment, and excellent cross-team communication skills

Preferred Qualifications:

  • Familiarity with MLOps practices, model serving/inferencing frameworks (e.g., vLLM, TGI, Triton Inference Server), and experience managing GPU clusters for AI/ML workloads
  • Proven track record working with large-scale distributed systems or high-concurrency environments
  • Hands-on experience in designing and building Internal Developer Platforms to enhance developer autonomy and productivity
  • Deep familiarity with Zero Trust architecture, automated security testing (DevSecOps), and implementing strict compliance frameworks
  • Prior experience acting as a Technical Lead, mentoring junior engineers, or managing DevOps teams
  • AIOps builder mindset: you see incidents, tickets, and manual runbooks as source material for the next automation, not as steady-state work
  • Paved-road philosophy: you think in golden paths and defaults, not policies; you'd rather make the right thing easy than write documentation scolding people for doing the wrong thing

How We Work With You

When you apply through Get A Job.ai, our talent team will review your qualifications and schedule an initial screening call. If there's a strong match, we'll prepare and submit your profile to our client. Throughout the interview process, we'll provide guidance and feedback to help you succeed.

Please apply exclusively through Get A Job.ai. Do not attempt to contact the client directly, as all applications must come through our recruiting process.

Pay

Pay information was not provided for this role. Compensation details will be discussed during the screening process.

Equal Employment Opportunity: Get A Job.ai is committed to providing equal employment opportunities to all candidates regardless of race, color, gender identity and/or expression, sexual orientation, marital status, religion, political opinion, nationality, ethnic background, disability, age, or any other protected characteristic.

Apply with Get A Job.ai

A recruiter will review your profile and submit you to the client. Do not contact the client directly.

More options

Apply with Get A Job.ai

Apply through Get A Job.ai. A recruiter will review your profile and submit you. Do not contact the client directly.

Apply through Get A Job.ai. A recruiter will review your profile and submit you.

Local insights for this role are preparing — this section updates automatically in a few seconds (or refresh).

Listing facts

  • Role Cloud Service Security Platform DevOps & Maintenance
  • Employer Get A Job.ai
  • Location United States
  • Type Full Time
  • Pay (from listing) Pay not listed
  • Posted September 9, 2026
  • Apply by October 9, 2026
  • Country United States
  • Overview Full job description on this page (781 words)

Facts above come from this job record on Get A Job.AI — not copied from third-party review sites.

Limited public data for this employer

We only show facts we can ground in public sources (Wikidata, O*NET, news/discussion links, or this listing). We do not invent Glassdoor-style ratings, salaries, or testimonials when data is thin. Use the listing facts, occupation context, and related openings below while we continue researching.

Explore related openings

Keep exploring on Get A Job.ai

Not quite the right fit? Your next opportunity is a click away.

Hiring instead? Post a job and reach candidates searching right now.

Cloud Service Security Platform DevOps & Ma… Get A Job.ai · United States