- Company: Get A Job.ai
- Location: USA
- Salary: Pay not listed
- Work type: Remote
Website Get A Job.ai
Represented by Get A Job.ai
About the Opportunity
We are representing a confidential SaaS organization seeking a Site Reliability Engineer 2 to join their global Platform SRE team. This is a hands-on role focused on building, operating, and scaling a multi-region cloud platform that serves thousands of customers across AWS, GCP, and Azure.
You'll work on production systems at enterprise scale, including multi-region Kubernetes clusters, service mesh architectures, and distributed gateway systems. This position is ideal for engineers who thrive on running production SaaS environments, automating operations, and continuously improving performance and resilience.
Responsibilities
- Operate and scale a global SaaS platform, ensuring reliability, availability, and performance across multiple regions and cloud providers
- Build, automate, and maintain Kubernetes-based infrastructure and deployment workflows using Terraform/Terragrunt, Helm, and ArgoCD
- Design, maintain, and optimize multi-region data and caching layers including PostgreSQL, Redis, ClickHouse, and Druid for high availability and low latency
- Operate and improve gateway and mesh environments supporting hybrid and distributed architectures
- Develop and maintain CI/CD pipelines and GitOps workflows to automate service delivery and ensure consistent infrastructure changes
- Enhance observability and incident response readiness through systems like Datadog, Prometheus, Grafana, and Thanos, defining and tracking SLOs
- Collaborate with development and security teams to ensure smooth operation of SaaS services in compliance with reliability, security, and regulatory standards
- Participate in a global 24/7 on-call rotation and drive continuous improvement of operational playbooks and postmortem practices
- Lead and contribute to scaling initiatives that improve elasticity, reliability, and cost-efficiency across the platform
What We're Looking For
Required Qualifications:
- BS in Computer Science or equivalent practical experience
- Proven experience managing SaaS or PaaS systems at enterprise scale (multi-region, multi-tenant, secure environments)
- Deep expertise in Kubernetes, including debugging cluster/networking issues and designing for fault tolerance and scalability
- Strong proficiency with Infrastructure as Code tools like Terraform or Terragrunt
- Experience with CI/CD pipelines and GitOps workflows (ArgoCD, Atlantis, Helm)
- Proficiency in one or more programming languages (Go, Python, Bash) for automation and tooling
- Solid understanding of Linux/Unix systems, networking (DNS, TLS/SSL, HTTP), load balancers, and distributed systems
- Experience working with API gateway and service mesh technologies
- Familiarity with streaming systems like Kafka and observability platforms (Datadog, Prometheus, Grafana)
- Experience working in a 24/7/365 production support environment
Bonus Qualifications:
- Hands-on experience with service connectivity technologies and gateway platforms
- Experience operating ClickHouse, Druid, or other time-series and analytics databases
- Experience managing PostgreSQL and Redis in multi-region configurations
- Working knowledge of AWS networking (PrivateLink, Transit Gateway, VPC Peering, Firewalls), Azure VNet, or GCP NCC
- Strong understanding of disaster recovery, resiliency testing, and compliance-driven reliability practices
How We Work With You
Our talent team at Get A Job.ai partners with leading technology companies to connect exceptional engineers with career-defining opportunities. When you apply through our platform, a dedicated recruiter will review your background and coordinate directly with you throughout the process. We handle the introduction to our client and guide you through each stage of interview and offer negotiations.
Please apply exclusively through Get A Job.ai — do not contact the client directly, as all candidates must be submitted through our formal partnership.
Pay
Compensation details will be discussed with qualified candidates during the screening process.
Equal Opportunity: Get A Job.ai is committed to building diverse and inclusive candidate pools. We encourage applications from individuals of all backgrounds and experiences.
Apply with Get A Job.ai
A recruiter will review your profile and submit you to the client. Do not contact the client directly.
Apply through Get A Job.ai. A recruiter will review your profile and submit you. Do not contact the client directly.
Apply through Get A Job.ai. A recruiter will review your profile and submit you.
Local insights for this role are preparing — this section updates automatically in a few seconds (or refresh).
Listing facts
- Role Site Reliability Engineer 2
- Employer Get A Job.ai
- Location USA · Remote-friendly
- Type Full Time
- Pay (from listing) Pay not listed
- Posted September 17, 2026
- Apply by October 17, 2026
- Country United States
- Overview Full job description on this page (548 words)
Facts above come from this job record on Get A Job.AI — not copied from third-party review sites.
Typical work in Senior Site Reliability Engineer
Independent occupational context from O*NET (U.S. public-domain labor data). This is about the occupation, not a rewrite of this employer's posting.
- Study product characteristics or customer requirements to determine validation objectives and standards.
- Analyze validation test data to determine whether systems or processes have met validation criteria or to identify root causes of production problems.
- Develop validation master plans, process flow diagrams, test cases, or standard operating procedures.
- Prepare detailed reports or design statements, based on results of validation and qualification tests or reviews of procedures and protocols.
- Maintain validation test equipment.
- Conduct validation or qualification tests of new or existing processes, equipment, or software in accordance with internal protocols or external standards.
Limited public data for this employer
We only show facts we can ground in public sources (Wikidata, O*NET, news/discussion links, or this listing). We do not invent Glassdoor-style ratings, salaries, or testimonials when data is thin. Use the listing facts, occupation context, and related openings below while we continue researching.
Explore related openings
Occupation family: Senior Site Reliability Engineer
Keep exploring on Get A Job.ai
Not quite the right fit? Your next opportunity is a click away.
- Browse all jobs
- More jobs by category
- Remote jobs you can do from anywhere
- Research typical pay for this role
- Set a job alert so new matches reach you first
- Upload your resume to apply faster
Hiring instead? Post a job and reach candidates searching right now.
