- Company: Get A Job.ai
- Location: USA
- Salary: Pay not listed
- Work type: Remote
Website Get A Job.ai
Represented by Get A Job.ai
About This Opportunity
Our talent team is representing a confidential AI infrastructure organization seeking a Senior Incident Manager to lead critical incident response across their data center operations. This role serves as the central command point during major service-impacting events, coordinating rapid resolution and driving operational resilience improvements.
This is a USA-based position requiring deep operational expertise in high-availability infrastructure, large-scale GPU compute environments, and cloud platforms, combined with exceptional leadership and crisis communication skills.
Responsibilities
- Lead end-to-end response for critical (SEV-1/SEV-2) incidents affecting AI infrastructure, GPU clusters, networking, storage, and data center operations
- Serve as Incident Commander during major outages, coordinating engineering, networking, facilities, and vendor teams
- Act as liaison between leadership and technical teams, providing clear status updates and incident summaries
- Own the complete incident lifecycle: technical triage assistance, escalation, coordination, resolution, and post-incident review
- Lead post-incident reviews (PIRs) and root cause analysis to identify systemic reliability gaps and implement corrective actions
- Work in an on-call rotation to respond to and coordinate incidents
- Maintain incident response documentation, operational playbooks, and runbooks
- Track and analyze incident metrics including MTTR, MTTD, and recurrence rates
- Improve incident response processes, escalation paths, and tooling in collaboration with technical support and engineering
- Provide executive-level incident reports and maintain operational health dashboards
- Collaborate across data center operations, infrastructure engineering, network engineering, platform reliability, security operations, and hardware vendors
What We're Looking For
Required Qualifications:
- 8+ years of experience in incident management, site reliability engineering, or infrastructure operations
- Proven track record managing incidents in large-scale distributed infrastructure environments
- Strong understanding of data center operations, GPU compute clusters, networking and storage infrastructure, and cloud/hybrid platforms
- Demonstrated ability to lead high-pressure incident response situations
- Experience with incident management frameworks (ITIL, SRE, or equivalent)
- Excellent communication and stakeholder management skills
- Hands-on experience with incident tracking and monitoring tools such as PagerDuty, ServiceNow, Jira, Datadog, Prometheus, or Grafana
Preferred Qualifications:
- Experience operating AI or HPC infrastructure
- Background in SRE, infrastructure engineering, or data center operations
- Familiarity with high-density GPU environments including NVIDIA clusters and InfiniBand networks
- Experience with hyperscale or colocation data center environments
- Knowledge of automation and incident response tooling
- Knowledge of and experience with Incident Command System (ICS)
- Experience leading and developing incident command programs from scratch
Key Competencies: Incident command and leadership, operational decision-making, cross-team coordination, root cause analysis, crisis communication, infrastructure reliability.
How We Work With You
Candidates apply directly through Get A Job.ai. Our recruiting team will conduct an initial screening to understand your background and confirm fit for the role. We then submit qualified candidates to our client for their review and interview process. Please do not attempt to contact the client directly, as all communications are managed through our team to ensure a smooth and professional process.
Pay
Compensation details will be discussed during the screening process based on your experience and qualifications.
Equal Opportunity: Get A Job.ai is committed to equal employment opportunity. We consider applicants without regard to race, color, religion, creed, national origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by law.
Apply with Get A Job.ai
A recruiter will review your profile and submit you to the client. Do not contact the client directly.
Apply through Get A Job.ai. A recruiter will review your profile and submit you. Do not contact the client directly.
Apply through Get A Job.ai. A recruiter will review your profile and submit you.
Local insights for this role are preparing — this section updates automatically in a few seconds (or refresh).
Listing facts
- Role Senior Incident Manager
- Employer Get A Job.ai
- Location USA · Remote-friendly
- Type Full Time
- Pay (from listing) Pay not listed
- Posted September 6, 2026
- Apply by October 6, 2026
- Country United States
- Overview Full job description on this page (518 words)
Facts above come from this job record on Get A Job.AI — not copied from third-party review sites.
Limited public data for this employer
We only show facts we can ground in public sources (Wikidata, O*NET, news/discussion links, or this listing). We do not invent Glassdoor-style ratings, salaries, or testimonials when data is thin. Use the listing facts, occupation context, and related openings below while we continue researching.
Explore related openings
Occupation family: Incident Manager
Keep exploring on Get A Job.ai
Not quite the right fit? Your next opportunity is a click away.
- Browse all jobs
- More jobs by category
- Remote jobs you can do from anywhere
- Research typical pay for this role
- Set a job alert so new matches reach you first
- Upload your resume to apply faster
Hiring instead? Post a job and reach candidates searching right now.
