- Company: Get A Job.ai
- Location: United States
- Salary: Pay not listed
Website Get A Job.ai
Represented by Get A Job.ai
About This Opportunity
We are representing a confidential healthcare technology organization in their search for a Senior Manager, Site Reliability Engineering. Our client operates a mission-critical SaaS platform that processes sensitive patient and commercial data, serving leading organizations in the biopharmaceutical sector. This is a player-coach role where you'll lead a high-impact SRE and DevOps team while remaining hands-on with infrastructure design and reliability engineering.
The platform runs on AWS and must maintain the highest standards of reliability, security, and compliance given the regulated nature of the data involved. You'll report to senior engineering leadership and partner closely with software engineering, product management, and security teams to ensure operational excellence across all production systems.
Responsibilities
As Senior Manager of Site Reliability Engineering, you will:
- Own reliability, availability, performance, and scalability of the AWS-hosted SaaS platform with accountability for SLA and SLO commitments
- Define and maintain Service Level Objectives and Service Level Indicators across production services; use error budgets to drive engineering prioritization
- Lead incident response and on-call operations including triage, coordination, stakeholder communication, and thorough post-incident reviews
- Drive proactive reliability culture through load testing, chaos engineering, and systematic failure mode analysis
- Architect and maintain AWS cloud-native infrastructure using infrastructure-as-code (AWS CDK, Terraform), ensuring environments are reproducible, auditable, and secure
- Oversee and continuously improve CI/CD pipelines (GitHub Actions) for rapid, safe delivery of application and infrastructure changes
- Manage core AWS services including ECS, EC2, Aurora RDS, DynamoDB, Lambda, S3, SQS, EventBridge, Cognito, Secrets Manager, and CloudFront
- Ensure robust observability across the stack using CloudWatch, Sentry, and related tooling for rapid detection and response
- Manage capacity planning, cost optimization, and cloud spend governance
- Ensure all infrastructure practices meet HIPAA, SOC 2, and HITRUST requirements for protected health information
- Partner with security teams on vulnerability management, infrastructure hardening, secrets management, and access control
- Maintain and regularly test disaster recovery and business continuity plans with defined RTO/RPO targets
- Lead, mentor, and grow a blended team of full-time SREs and offshore contractors, fostering ownership and continuous improvement
- Manage distributed team dynamics with clear communication rhythms, documentation standards, and handoff protocols
- Build and maintain healthy on-call rotation with appropriate tooling, runbooks, and escalation paths
- Collaborate with software engineering to embed reliability practices into the development lifecycle
- Champion automation-first thinking to eliminate toil through tooling and process improvement
What We're Looking For
Required qualifications:
- 7+ years of experience in SRE, DevOps, or infrastructure engineering
- At least 3 years in people management or team lead capacity
- Deep, hands-on AWS expertise with ability to architect, operate, and optimize cloud-native workloads at the service level
- Strong infrastructure-as-code skills with AWS CDK, Terraform, or equivalent tools
- Demonstrated experience owning SLOs, incident management, and on-call operations in commercial SaaS environments
- Experience managing CI/CD pipelines and developer productivity tooling with understanding of deployment safety (canary releases, feature flags, rollback strategies)
- Working knowledge of security and compliance requirements for regulated data environments (HIPAA, SOC 2, HITRUST)
- Proven ability to lead and develop teams, including working effectively with offshore or distributed contractors across time zones
- Strong written and verbal communication skills with ability to explain infrastructure risk and trade-offs to technical and non-technical audiences
- Comfort operating in fast-paced, high-growth startup environments where priorities evolve and initiative is expected
- Role is primarily remote with occasional travel requirements
Preferred qualifications:
- Experience in healthcare technology or digital health with direct exposure to HIPAA-regulated protected health information
- Familiarity with TypeScript, NestJS, PostgreSQL, and React
- Experience with chaos engineering practices and tools
- Experience leveraging AI tools to improve operational workflows, accelerate runbook development, or automate incident triage
- Bachelor's degree in Computer Science, Engineering, or related discipline, or equivalent practical experience
Technology stack you'll work with:
- Cloud: AWS (ECS, EC2, Aurora RDS, DynamoDB, Lambda, S3, SQS, EventBridge, Cognito, Secrets Manager, CloudFront, WAF)
- Infrastructure as Code: AWS CDK (primary); Terraform/OpenTofu familiarity a plus
- CI/CD: GitHub Actions
- Application Platform: NestJS/TypeScript, React/TypeScript, PostgreSQL, Turborepo, pnpm
- Data Platform: AWS Glue, Lake Formation, PySpark, Kinesis, Python-based ELT pipelines
- Observability: CloudWatch, Sentry, OpenFeature, Tableau
- Languages: TypeScript, Python, SQL; Golang familiarity a plus
How We Work With You
When you apply through Get A Job.ai, our talent team will carefully review your background against the role requirements. Qualified candidates will be contacted by one of our recruiters for an initial screening conversation to discuss your experience, technical expertise, and career goals.
If there's a strong mutual fit, we'll prepare and submit your profile to our client for consideration. Throughout the process, we'll keep you informed of progress and provide guidance on next steps. Our client conducts thorough technical interviews including multiple video conversations with engineering leadership and team members.
Important: Please apply exclusively through Get A Job.ai. Do not attempt to contact the client directly, as this may disqualify your application. All communication regarding this opportunity should go through our recruiting team.
Pay
Pay via Get A Job.ai: $155,000 to $170,000
Our client also offers a competitive benefits package including unlimited PTO, stock options, and comprehensive health benefits. For employees within reasonable driving distance of regional hubs, the organization hosts in-person gatherings approximately every other month.
Get A Job.ai is an equal opportunity employer. We are committed to building diverse, inclusive teams and encourage applications from candidates of all backgrounds. We provide reasonable accommodations to applicants with disabilities throughout our recruiting process.
Apply with Get A Job.ai
A recruiter will review your profile and submit you to the client. Do not contact the client directly.
Apply through Get A Job.ai. A recruiter will review your profile and submit you. Do not contact the client directly.
Apply through Get A Job.ai. A recruiter will review your profile and submit you.
Terms used in this posting
- PTO
- Paid Time Off — vacation, personal, or sick days you can take while still being paid.
- on-call
- You may be required to be reachable and available to work outside normal scheduled hours, typically for a set rotation.
Explore Get A Job.ai online
Working in United States
The United States of America (USA), also known as the United States (U.S.) or America, is a country primarily located in North America. It is a federal republic consisting of 50 states and a federal capital district, Washington, D.C. The 48 contiguous states border Canada to the north and Mexico to the south, with the semi-exclave of Alaska in the northwest and the archipelago of Hawaii in the Pacific Ocean. The United States also asserts sovereignty over five major island territories and various uninhabited islands in Oceania and the Caribbean. It is a megadiverse country, with the world's th
🇺🇸 Relocation safety for US: Exercise Normal Caution — via Warnely, CC BY 4.0
National unemployment rate in US: 4.2% — via World Bank
National job openings rate: 4.4% — via BLS JOLTS
Private-sector wage growth (year over year): 3.1% — via FRED
National quits rate: 2.0% — via FRED (BLS JOLTS)
Weekly initial unemployment claims: 196,000 — via FRED
GDP per capita in US: $90,027 — via World Bank
Consumer price inflation in US: 2.9% (annual) — via World Bank
Real GDP growth in US: 2.2% (annual) — via World Bank
Average hours worked per year in US: 1,800 — via OECD
Job details above are provided by the employer/source. The sections on this page are compiled from public data sources with AI assistance.
Accommodations: if you need a workplace accommodation to apply for or perform this job, see ADA.gov or EEOC.gov for guidance on your rights and how to request one.
Listing facts
- Role Senior Manager, Site Reliability Engineering
- Employer Get A Job.ai
- Location United States
- Type Full Time
- Pay (from listing) Pay not listed
- Posted September 11, 2026
- Apply by October 12, 2026
- Country United States
- Overview Full job description on this page (879 words)
Facts above come from this job record on Get A Job.AI — not copied from third-party review sites.
Limited public data for this employer
We only show facts we can ground in public sources (Wikidata, O*NET, news/discussion links, or this listing). We do not invent Glassdoor-style ratings, salaries, or testimonials when data is thin. Use the listing facts, occupation context, and related openings below while we continue researching.
Explore related openings
Keep exploring on Get A Job.ai
Not quite the right fit? Your next opportunity is a click away.
- Browse all jobs
- More jobs by category
- Remote jobs you can do from anywhere
- Research typical pay for this role
- Set a job alert so new matches reach you first
- Upload your resume to apply faster
Hiring instead? Post a job and reach candidates searching right now.
