- Company: Get A Job.ai
- Location: United States
- Salary: Pay not listed
Website Get A Job.ai
Represented by Get A Job.ai
About This Opportunity
Our talent team is representing a confidential managed services organization in their search for a Senior Automation & Observability Engineer. This is a remote role based in the United States, supporting enterprise infrastructure, applications, and telemetry ecosystems across a large-scale IT environment.
You will work with observability platforms, monitoring technologies, and automation frameworks to ensure high availability and performance of business-critical systems. This role involves close collaboration with infrastructure, cloud, network, application, and service delivery teams.
Responsibilities
Monitoring & Observability:
- Design, implement, and maintain enterprise monitoring and observability solutions
- Develop and maintain dashboards, alerts, and visualizations using Grafana
- Monitor infrastructure, applications, middleware, and IoT services using IBM Instana, Grafana, SolarWinds, and related tools
- Configure and manage data collection using Telegraf, Prometheus, and monitoring agents
- Analyze metrics, logs, traces, events, and telemetry data to identify performance bottlenecks
- Support SLO, SLA, and operational health monitoring initiatives
- Perform Root Cause Analysis and troubleshooting for infrastructure and application issues
Enterprise Logging & Telemetry:
- Support onboarding and operational management of enterprise applications and FOAK services
- Configure, validate, and troubleshoot Enterprise Logging & Telemetry integrations across infrastructure, middleware, applications, and cloud platforms
- Monitor telemetry pipelines, log ingestion, event correlation, and data quality
- Collaborate with engineering teams to improve telemetry standards and monitoring effectiveness
Infrastructure & Platform Monitoring:
- Monitor Linux Servers, Windows Servers, VMware Infrastructure, Citrix VDI Platforms, DNS Services, Proxy Services, Middleware Platforms, and Integration Services
- Investigate performance issues, recurring alerts, and infrastructure anomalies
- Monitor capacity, availability, CPU, memory, storage, and service health metrics
- Support platform upgrades, maintenance, and operational readiness reviews
Database & Data Management:
- Configure and maintain InfluxDB time-series databases
- Manage data retention policies, performance tuning, and capacity planning
- Develop operational dashboards and reports for infrastructure and application performance insights
Event & Incident Management:
- Monitor operational alerts, events, notifications, and incidents from enterprise monitoring platforms
- Acknowledge, investigate, troubleshoot, and resolve assigned incidents
- Coordinate with cross-functional teams during incident resolution
- Participate in major incident bridges, DR exercises, and 24x7 operations support activities
- Follow escalation procedures, SOPs, operational runbooks, and ITIL processes
Instana & APM Operations:
- Administer and support IBM Instana monitoring environments
- Monitor application, API, middleware, and microservices performance using Instana
- Validate Instana agent health following server patching and maintenance activities
- Configure alerts, baselines, and performance thresholds
Automation & Scripting:
- Develop automation solutions using Python, PowerShell, Shell Scripting (Bash), and VBScript
- Automate operational tasks, monitoring deployments, and remediation workflows
- Build reusable automation tools to improve operational efficiency
- Integrate monitoring platforms with enterprise automation frameworks
Configuration Management & Infrastructure Automation:
- Implement Infrastructure as Code (IaC) and automation using Ansible and Puppet
- Automate server provisioning and configuration management
- Automate monitoring agent deployment and onboarding
- Maintain automation playbooks and deployment pipelines
Knowledge Management:
- Maintain SOPs, runbooks, monitoring procedures, and escalation matrices
- Participate in knowledge transfer sessions, service onboarding, and operational readiness reviews
- Support service transition, migration, and continuous improvement initiatives
What We're Looking For
Required Skills:
- 7 to 10+ years of experience in Monitoring, Observability, Infrastructure Operations, SRE, or Platform Engineering
- Hands-on experience with Grafana, IBM Instana, SolarWinds, Telegraf, Prometheus, and InfluxDB
- Experience supporting large-scale enterprise environments and 24x7 operations
- Expertise in VMware, Linux Administration, Windows Server, Citrix VDI, DNS Services, Proxy Services, and Middleware Technologies
- Strong scripting skills in Python, PowerShell, Shell Scripting (Linux), and VBScript
- Experience with automation tools including Ansible and Puppet
- FOAK application support and Enterprise Logging & Telemetry (ELT) experience
- Experience with ServiceNow, Incident Management, Problem Management, Change Management, and ITIL Framework
- Strong troubleshooting, RCA, and operational support skills
- Experience integrating observability platforms with enterprise automation solutions
Preferred Skills:
- Docker and Kubernetes
- Cloud platforms: AWS, Microsoft Azure, Google Cloud Platform (GCP)
- CI/CD tools: Jenkins, GitHub Actions, GitLab CI/CD
- REST APIs and Microservices Monitoring
- DevOps & SRE Practices
How We Work With You
When you apply through Get A Job.ai, our recruiting team will review your qualifications and experience against our client's requirements. If we determine there's a strong match, we'll conduct an initial screening conversation to discuss your background, career goals, and the role in detail.
After our screening, we will submit qualified candidates to our client for consideration. Throughout the interview process, we remain your advocate and main point of contact. Please do not contact the client organization directly—all communication should flow through Get A Job.ai to ensure a smooth and professional process.
Pay
Pay details will be discussed during the screening process based on your experience and qualifications.
Equal Employment Opportunity: Get A Job.ai is committed to providing equal employment opportunities to all applicants. We evaluate candidates based on qualifications and merit without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, veteran status, or other legally protected characteristics.
Apply with Get A Job.ai
A recruiter will review your profile and submit you to the client. Do not contact the client directly.
Apply through Get A Job.ai. A recruiter will review your profile and submit you. Do not contact the client directly.
Apply through Get A Job.ai. A recruiter will review your profile and submit you.
Explore Get A Job.ai online
Working in United States
The United States of America (USA), also known as the United States (U.S.) or America, is a country primarily located in North America. It is a federal republic consisting of 50 states and a federal capital district, Washington, D.C. The 48 contiguous states border Canada to the north and Mexico to the south, with the semi-exclave of Alaska in the northwest and the archipelago of Hawaii in the Pacific Ocean. The United States also asserts sovereignty over five major island territories and various uninhabited islands in Oceania and the Caribbean. It is a megadiverse country, with the world's th
🇺🇸 Relocation safety for US: Exercise Normal Caution — via Warnely, CC BY 4.0
National unemployment rate in US: 4.2% — via World Bank
National job openings rate: 4.4% — via BLS JOLTS
Private-sector wage growth (year over year): 3.1% — via FRED
National quits rate: 2.0% — via FRED (BLS JOLTS)
Weekly initial unemployment claims: 206,000 — via FRED
GDP per capita in US: $90,027 — via World Bank
Consumer price inflation in US: 2.9% (annual) — via World Bank
Real GDP growth in US: 2.2% (annual) — via World Bank
Average hours worked per year in US: 1,800 — via OECD
Job details above are provided by the employer/source. The sections on this page are compiled from public data sources with AI assistance.
Accommodations: if you need a workplace accommodation to apply for or perform this job, see ADA.gov or EEOC.gov for guidance on your rights and how to request one.
Listing facts
- Role Senior Automation & Observability Engineer
- Employer Get A Job.ai
- Location United States
- Type Full Time
- Pay (from listing) Pay not listed
- Posted September 19, 2026
- Apply by October 19, 2026
- Country United States
- Overview Full job description on this page (779 words)
Facts above come from this job record on Get A Job.AI — not copied from third-party review sites.
Limited public data for this employer
We only show facts we can ground in public sources (Wikidata, O*NET, news/discussion links, or this listing). We do not invent Glassdoor-style ratings, salaries, or testimonials when data is thin. Use the listing facts, occupation context, and related openings below while we continue researching.
Explore related openings
Keep exploring on Get A Job.ai
Not quite the right fit? Your next opportunity is a click away.
- Browse all jobs
- More jobs by category
- Remote jobs you can do from anywhere
- Research typical pay for this role
- Set a job alert so new matches reach you first
- Upload your resume to apply faster
Hiring instead? Post a job and reach candidates searching right now.
