- Company: Get A Job.ai
- Salary: Pay not listed
- Work type: Remote
Website Get A Job.ai
Represented by Get A Job.ai
About This Opportunity
We are representing a confidential enterprise software organization that is building the future of intelligent process orchestration. Our client is a global, fully remote company trusted by hundreds of major enterprises worldwide and recognized as a leader in their space. They are seeking a Senior Site Reliability Engineer to join their infrastructure team and help scale their multi-cloud Kubernetes platform.
This is a remote position where you'll design and maintain critical infrastructure, improve observability across the stack, and work directly with product and engineering teams to deliver reliable, scalable systems. You'll own the systems that power their platform, drive meaningful automation, and mentor other engineers along the way.
Responsibilities
- Design and maintain Kubernetes-based, multi-cloud platform architecture ensuring high availability, scalability, and fault tolerance
- Establish configuration best practices and network services that development teams depend on
- Implement and improve monitoring and alerting tools using platforms like Prometheus and Grafana to provide visibility into system health and performance
- Participate in on-call rotations and provide expert-level incident response and troubleshooting
- Create runbooks and automation that transform complex problems into manageable, repeatable processes
- Collaborate cross-functionally with product engineering, product management, and support teams to define and deliver infrastructure features
- Identify repetitive operational work and automate it, sharing knowledge with teammates to raise the overall quality bar
- Mentor less experienced engineers on complex infrastructure challenges and help them develop their technical skills
- Conduct root cause analysis on production incidents and implement preventive measures
What We're Looking For
Required qualifications:
- Deep hands-on experience building, deploying, and maintaining production Kubernetes clusters at scale
- Strong understanding of Kubernetes workload management, networking, and storage
- Expertise with infrastructure as code tools like Terraform or similar IaC platforms
- Demonstrated experience implementing monitoring and observability solutions with tools like Prometheus, Grafana, or equivalent
- Proven track record providing 3rd-level support and incident response in production environments
- Strong communication skills for explaining complex technical issues to stakeholders under pressure
- Passion for automation and continuously improving system reliability and maintainability
- Responsible use of AI tools for research, code review, documentation, and automation while maintaining human accountability
Preferred qualifications:
- Experience with AWS (EKS), Google Cloud Platform (GKE), or other managed Kubernetes services
- Familiarity with ArgoCD or GitOps workflows for declarative infrastructure management
- Proficiency in Python, Go, or similar languages for building automation tools
- Experience defining service level objectives (SLOs) and implementing effective alerting frameworks
How We Work With You
When you apply through Get A Job.ai, our talent team will review your background and schedule an initial screening call to understand your experience and career goals. If there's a strong match, we'll submit your profile to our client for consideration. Throughout the interview process, we'll provide guidance and feedback to help you succeed.
Please note: All applications must go through Get A Job.ai. Do not attempt to contact the client directly, as this may disqualify your candidacy.
Pay
Compensation details will be discussed during the screening process and are competitive with market rates for senior SRE positions. Final offers depend on skills, experience, and location.
Get A Job.ai is an equal opportunity recruiter. We welcome applications from all qualified candidates regardless of gender, race, ethnicity, religion, sexual orientation, age, disability, or any other protected characteristic. We are committed to presenting diverse talent to our clients.
Apply with Get A Job.ai
A recruiter will review your profile and submit you to the client. Do not contact the client directly.
Apply through Get A Job.ai. A recruiter will review your profile and submit you. Do not contact the client directly.
Apply through Get A Job.ai. A recruiter will review your profile and submit you.
Terms used in this posting
- on-call
- You may be required to be reachable and available to work outside normal scheduled hours, typically for a set rotation.
Explore Get A Job.ai online
About this role & career path
Traits that fit this role
- Achievement Orientation
- Intellectual Curiosity
- Cautiousness
- Integrity
- Attention to Detail
Source: O*NET Work Styles (Distinctiveness Rank).
Typical preparation needed: Job Zone 4: Considerable Preparation Needed. Most of these occupations require a four-year bachelor's degree, but some do not. — via O*NET
Industry news
- Inaugural Architecture & Engineering Industry Day - portseattle.org
- Anie: the industry fears a void following the NRRP: ‘We now need continuity in investment’ - Il Sole 24 ORE
- Mechanical engineering scholarship delivers immediate impact - Texas A&M University
Source: O*NET (public-domain bulk data)
Salary & compensation
Workers in Architecture & Engineering occupations earn a national median of $95,541 — via US Census ACS / Data USA
Market context
- Similar listings in Remote 1
Job details above are provided by the employer/source. The sections on this page are compiled from public data sources with AI assistance.
Accommodations: if you need a workplace accommodation to apply for or perform this job, see ADA.gov or EEOC.gov for guidance on your rights and how to request one.
Listing facts
- Role Senior Site Reliability Engineer
- Employer Get A Job.ai
- Location Remote · Remote-friendly
- Type Full Time
- Pay (from listing) Pay not listed
- Posted September 23, 2026
- Apply by October 23, 2026
- Overview Full job description on this page (545 words)
Facts above come from this job record on Get A Job.AI — not copied from third-party review sites.
Typical work in Senior Site Reliability Engineer
Independent occupational context from O*NET (U.S. public-domain labor data). This is about the occupation, not a rewrite of this employer's posting.
- Study product characteristics or customer requirements to determine validation objectives and standards.
- Analyze validation test data to determine whether systems or processes have met validation criteria or to identify root causes of production problems.
- Develop validation master plans, process flow diagrams, test cases, or standard operating procedures.
- Prepare detailed reports or design statements, based on results of validation and qualification tests or reviews of procedures and protocols.
- Maintain validation test equipment.
- Conduct validation or qualification tests of new or existing processes, equipment, or software in accordance with internal protocols or external standards.
Limited public data for this employer
We only show facts we can ground in public sources (Wikidata, O*NET, news/discussion links, or this listing). We do not invent Glassdoor-style ratings, salaries, or testimonials when data is thin. Use the listing facts, occupation context, and related openings below while we continue researching.
Explore related openings
Occupation family: Senior Site Reliability Engineer
Keep exploring on Get A Job.ai
Not quite the right fit? Your next opportunity is a click away.
- Browse all jobs
- More jobs by category
- Remote jobs you can do from anywhere
- Research typical pay for this role
- Set a job alert so new matches reach you first
- Upload your resume to apply faster
Hiring instead? Post a job and reach candidates searching right now.
