Loading...

Senior Manager, Site Reliability Engineering

  • Company: Get A Job.ai
  • Location: United States
  • Salary: Pay not listed

Website Get A Job.ai

Represented by Get A Job.ai

About This Opportunity

We are representing a confidential healthcare technology organization in their search for a Senior Manager, Site Reliability Engineering. Our client operates a mission-critical SaaS platform that processes sensitive patient and commercial data, serving leading organizations in the biopharmaceutical sector. This is a player-coach role where you'll lead a high-impact SRE and DevOps team while remaining hands-on with infrastructure design and reliability engineering.

The platform runs on AWS and must maintain the highest standards of reliability, security, and compliance given the regulated nature of the data involved. You'll report to senior engineering leadership and partner closely with software engineering, product management, and security teams to ensure operational excellence across all production systems.

Responsibilities

As Senior Manager of Site Reliability Engineering, you will:

  • Own reliability, availability, performance, and scalability of the AWS-hosted SaaS platform with accountability for SLA and SLO commitments
  • Define and maintain Service Level Objectives and Service Level Indicators across production services; use error budgets to drive engineering prioritization
  • Lead incident response and on-call operations including triage, coordination, stakeholder communication, and thorough post-incident reviews
  • Drive proactive reliability culture through load testing, chaos engineering, and systematic failure mode analysis
  • Architect and maintain AWS cloud-native infrastructure using infrastructure-as-code (AWS CDK, Terraform), ensuring environments are reproducible, auditable, and secure
  • Oversee and continuously improve CI/CD pipelines (GitHub Actions) for rapid, safe delivery of application and infrastructure changes
  • Manage core AWS services including ECS, EC2, Aurora RDS, DynamoDB, Lambda, S3, SQS, EventBridge, Cognito, Secrets Manager, and CloudFront
  • Ensure robust observability across the stack using CloudWatch, Sentry, and related tooling for rapid detection and response
  • Manage capacity planning, cost optimization, and cloud spend governance
  • Ensure all infrastructure practices meet HIPAA, SOC 2, and HITRUST requirements for protected health information
  • Partner with security teams on vulnerability management, infrastructure hardening, secrets management, and access control
  • Maintain and regularly test disaster recovery and business continuity plans with defined RTO/RPO targets
  • Lead, mentor, and grow a blended team of full-time SREs and offshore contractors, fostering ownership and continuous improvement
  • Manage distributed team dynamics with clear communication rhythms, documentation standards, and handoff protocols
  • Build and maintain healthy on-call rotation with appropriate tooling, runbooks, and escalation paths
  • Collaborate with software engineering to embed reliability practices into the development lifecycle
  • Champion automation-first thinking to eliminate toil through tooling and process improvement

What We're Looking For

Required qualifications:

  • 7+ years of experience in SRE, DevOps, or infrastructure engineering
  • At least 3 years in people management or team lead capacity
  • Deep, hands-on AWS expertise with ability to architect, operate, and optimize cloud-native workloads at the service level
  • Strong infrastructure-as-code skills with AWS CDK, Terraform, or equivalent tools
  • Demonstrated experience owning SLOs, incident management, and on-call operations in commercial SaaS environments
  • Experience managing CI/CD pipelines and developer productivity tooling with understanding of deployment safety (canary releases, feature flags, rollback strategies)
  • Working knowledge of security and compliance requirements for regulated data environments (HIPAA, SOC 2, HITRUST)
  • Proven ability to lead and develop teams, including working effectively with offshore or distributed contractors across time zones
  • Strong written and verbal communication skills with ability to explain infrastructure risk and trade-offs to technical and non-technical audiences
  • Comfort operating in fast-paced, high-growth startup environments where priorities evolve and initiative is expected
  • Role is primarily remote with occasional travel requirements

Preferred qualifications:

  • Experience in healthcare technology or digital health with direct exposure to HIPAA-regulated protected health information
  • Familiarity with TypeScript, NestJS, PostgreSQL, and React
  • Experience with chaos engineering practices and tools
  • Experience leveraging AI tools to improve operational workflows, accelerate runbook development, or automate incident triage
  • Bachelor's degree in Computer Science, Engineering, or related discipline, or equivalent practical experience

Technology stack you'll work with:

  • Cloud: AWS (ECS, EC2, Aurora RDS, DynamoDB, Lambda, S3, SQS, EventBridge, Cognito, Secrets Manager, CloudFront, WAF)
  • Infrastructure as Code: AWS CDK (primary); Terraform/OpenTofu familiarity a plus
  • CI/CD: GitHub Actions
  • Application Platform: NestJS/TypeScript, React/TypeScript, PostgreSQL, Turborepo, pnpm
  • Data Platform: AWS Glue, Lake Formation, PySpark, Kinesis, Python-based ELT pipelines
  • Observability: CloudWatch, Sentry, OpenFeature, Tableau
  • Languages: TypeScript, Python, SQL; Golang familiarity a plus

How We Work With You

When you apply through Get A Job.ai, our talent team will carefully review your background against the role requirements. Qualified candidates will be contacted by one of our recruiters for an initial screening conversation to discuss your experience, technical expertise, and career goals.

If there's a strong mutual fit, we'll prepare and submit your profile to our client for consideration. Throughout the process, we'll keep you informed of progress and provide guidance on next steps. Our client conducts thorough technical interviews including multiple video conversations with engineering leadership and team members.

Important: Please apply exclusively through Get A Job.ai. Do not attempt to contact the client directly, as this may disqualify your application. All communication regarding this opportunity should go through our recruiting team.

Pay

Pay via Get A Job.ai: $155,000 to $170,000

Our client also offers a competitive benefits package including unlimited PTO, stock options, and comprehensive health benefits. For employees within reasonable driving distance of regional hubs, the organization hosts in-person gatherings approximately every other month.

Get A Job.ai is an equal opportunity employer. We are committed to building diverse, inclusive teams and encourage applications from candidates of all backgrounds. We provide reasonable accommodations to applicants with disabilities throughout our recruiting process.

Apply with Get A Job.ai

A recruiter will review your profile and submit you to the client. Do not contact the client directly.

More options

Apply with Get A Job.ai

Apply through Get A Job.ai. A recruiter will review your profile and submit you. Do not contact the client directly.

Apply through Get A Job.ai. A recruiter will review your profile and submit you.

Terms used in this posting

PTO
Paid Time Off — vacation, personal, or sick days you can take while still being paid.
on-call
You may be required to be reachable and available to work outside normal scheduled hours, typically for a set rotation.

Working in United States

The United States of America (USA), also known as the United States (U.S.) or America, is a country primarily located in North America. It is a federal republic consisting of 50 states and a federal capital district, Washington, D.C. The 48 contiguous states border Canada to the north and Mexico to the south, with the semi-exclave of Alaska in the northwest and the archipelago of Hawaii in the Pacific Ocean. The United States also asserts sovereignty over five major island territories and various uninhabited islands in Oceania and the Caribbean. It is a megadiverse country, with the world's th

🇺🇸 Relocation safety for US: Exercise Normal Cautionvia Warnely, CC BY 4.0

National unemployment rate in US: 4.2%via World Bank

National job openings rate: 4.4%via BLS JOLTS

Private-sector wage growth (year over year): 3.1%via FRED

National quits rate: 2.0%via FRED (BLS JOLTS)

Weekly initial unemployment claims: 196,000via FRED

GDP per capita in US: $90,027via World Bank

Consumer price inflation in US: 2.9% (annual) — via World Bank

Real GDP growth in US: 2.2% (annual) — via World Bank

Average hours worked per year in US: 1,800via OECD

    Job details above are provided by the employer/source. The sections on this page are compiled from public data sources with AI assistance.

    Accommodations: if you need a workplace accommodation to apply for or perform this job, see ADA.gov or EEOC.gov for guidance on your rights and how to request one.

    Add application deadline to calendar

    Listing facts

    • Role Senior Manager, Site Reliability Engineering
    • Employer Get A Job.ai
    • Location United States
    • Type Full Time
    • Pay (from listing) Pay not listed
    • Posted September 11, 2026
    • Apply by October 12, 2026
    • Country United States
    • Overview Full job description on this page (879 words)

    Facts above come from this job record on Get A Job.AI — not copied from third-party review sites.

    Limited public data for this employer

    We only show facts we can ground in public sources (Wikidata, O*NET, news/discussion links, or this listing). We do not invent Glassdoor-style ratings, salaries, or testimonials when data is thin. Use the listing facts, occupation context, and related openings below while we continue researching.

    Explore related openings

    Keep exploring on Get A Job.ai

    Not quite the right fit? Your next opportunity is a click away.

    Hiring instead? Post a job and reach candidates searching right now.

    Senior Manager, Site Reliability Engineering Get A Job.ai · United States