Loading...

Senior Incident Manager

  • Company: Get A Job.ai
  • Location: USA
  • Salary: Pay not listed
  • Work type: Remote

Website Get A Job.ai

Represented by Get A Job.ai

About This Opportunity

Our talent team is representing a confidential AI infrastructure organization seeking a Senior Incident Manager to lead critical incident response across their data center operations. This role serves as the central command point during major service-impacting events, coordinating rapid resolution and driving operational resilience improvements.

This is a USA-based position requiring deep operational expertise in high-availability infrastructure, large-scale GPU compute environments, and cloud platforms, combined with exceptional leadership and crisis communication skills.

Responsibilities

  • Lead end-to-end response for critical (SEV-1/SEV-2) incidents affecting AI infrastructure, GPU clusters, networking, storage, and data center operations
  • Serve as Incident Commander during major outages, coordinating engineering, networking, facilities, and vendor teams
  • Act as liaison between leadership and technical teams, providing clear status updates and incident summaries
  • Own the complete incident lifecycle: technical triage assistance, escalation, coordination, resolution, and post-incident review
  • Lead post-incident reviews (PIRs) and root cause analysis to identify systemic reliability gaps and implement corrective actions
  • Work in an on-call rotation to respond to and coordinate incidents
  • Maintain incident response documentation, operational playbooks, and runbooks
  • Track and analyze incident metrics including MTTR, MTTD, and recurrence rates
  • Improve incident response processes, escalation paths, and tooling in collaboration with technical support and engineering
  • Provide executive-level incident reports and maintain operational health dashboards
  • Collaborate across data center operations, infrastructure engineering, network engineering, platform reliability, security operations, and hardware vendors

What We're Looking For

Required Qualifications:

  • 8+ years of experience in incident management, site reliability engineering, or infrastructure operations
  • Proven track record managing incidents in large-scale distributed infrastructure environments
  • Strong understanding of data center operations, GPU compute clusters, networking and storage infrastructure, and cloud/hybrid platforms
  • Demonstrated ability to lead high-pressure incident response situations
  • Experience with incident management frameworks (ITIL, SRE, or equivalent)
  • Excellent communication and stakeholder management skills
  • Hands-on experience with incident tracking and monitoring tools such as PagerDuty, ServiceNow, Jira, Datadog, Prometheus, or Grafana

Preferred Qualifications:

  • Experience operating AI or HPC infrastructure
  • Background in SRE, infrastructure engineering, or data center operations
  • Familiarity with high-density GPU environments including NVIDIA clusters and InfiniBand networks
  • Experience with hyperscale or colocation data center environments
  • Knowledge of automation and incident response tooling
  • Knowledge of and experience with Incident Command System (ICS)
  • Experience leading and developing incident command programs from scratch

Key Competencies: Incident command and leadership, operational decision-making, cross-team coordination, root cause analysis, crisis communication, infrastructure reliability.

How We Work With You

Candidates apply directly through Get A Job.ai. Our recruiting team will conduct an initial screening to understand your background and confirm fit for the role. We then submit qualified candidates to our client for their review and interview process. Please do not attempt to contact the client directly, as all communications are managed through our team to ensure a smooth and professional process.

Pay

Compensation details will be discussed during the screening process based on your experience and qualifications.

Equal Opportunity: Get A Job.ai is committed to equal employment opportunity. We consider applicants without regard to race, color, religion, creed, national origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by law.

Apply with Get A Job.ai

A recruiter will review your profile and submit you to the client. Do not contact the client directly.

More options

Apply with Get A Job.ai

Apply through Get A Job.ai. A recruiter will review your profile and submit you. Do not contact the client directly.

Apply through Get A Job.ai. A recruiter will review your profile and submit you.

Local insights for this role are preparing — this section updates automatically in a few seconds (or refresh).

Listing facts

  • Role Senior Incident Manager
  • Employer Get A Job.ai
  • Location USA · Remote-friendly
  • Type Full Time
  • Pay (from listing) Pay not listed
  • Posted September 6, 2026
  • Apply by October 6, 2026
  • Country United States
  • Overview Full job description on this page (518 words)

Facts above come from this job record on Get A Job.AI — not copied from third-party review sites.

Limited public data for this employer

We only show facts we can ground in public sources (Wikidata, O*NET, news/discussion links, or this listing). We do not invent Glassdoor-style ratings, salaries, or testimonials when data is thin. Use the listing facts, occupation context, and related openings below while we continue researching.

Explore related openings

Occupation family: Incident Manager

Keep exploring on Get A Job.ai

Not quite the right fit? Your next opportunity is a click away.

Hiring instead? Post a job and reach candidates searching right now.

Senior Incident Manager Get A Job.ai · USA