Loading...

Senior Automation & Observability Engineer

  • Company: Get A Job.ai
  • Location: United States
  • Salary: Pay not listed

Website Get A Job.ai

Represented by Get A Job.ai

About This Opportunity

Our talent team is representing a confidential managed services organization in their search for a Senior Automation & Observability Engineer. This is a remote role based in the United States, supporting enterprise infrastructure, applications, and telemetry ecosystems across a large-scale IT environment.

You will work with observability platforms, monitoring technologies, and automation frameworks to ensure high availability and performance of business-critical systems. This role involves close collaboration with infrastructure, cloud, network, application, and service delivery teams.

Responsibilities

Monitoring & Observability:

  • Design, implement, and maintain enterprise monitoring and observability solutions
  • Develop and maintain dashboards, alerts, and visualizations using Grafana
  • Monitor infrastructure, applications, middleware, and IoT services using IBM Instana, Grafana, SolarWinds, and related tools
  • Configure and manage data collection using Telegraf, Prometheus, and monitoring agents
  • Analyze metrics, logs, traces, events, and telemetry data to identify performance bottlenecks
  • Support SLO, SLA, and operational health monitoring initiatives
  • Perform Root Cause Analysis and troubleshooting for infrastructure and application issues

Enterprise Logging & Telemetry:

  • Support onboarding and operational management of enterprise applications and FOAK services
  • Configure, validate, and troubleshoot Enterprise Logging & Telemetry integrations across infrastructure, middleware, applications, and cloud platforms
  • Monitor telemetry pipelines, log ingestion, event correlation, and data quality
  • Collaborate with engineering teams to improve telemetry standards and monitoring effectiveness

Infrastructure & Platform Monitoring:

  • Monitor Linux Servers, Windows Servers, VMware Infrastructure, Citrix VDI Platforms, DNS Services, Proxy Services, Middleware Platforms, and Integration Services
  • Investigate performance issues, recurring alerts, and infrastructure anomalies
  • Monitor capacity, availability, CPU, memory, storage, and service health metrics
  • Support platform upgrades, maintenance, and operational readiness reviews

Database & Data Management:

  • Configure and maintain InfluxDB time-series databases
  • Manage data retention policies, performance tuning, and capacity planning
  • Develop operational dashboards and reports for infrastructure and application performance insights

Event & Incident Management:

  • Monitor operational alerts, events, notifications, and incidents from enterprise monitoring platforms
  • Acknowledge, investigate, troubleshoot, and resolve assigned incidents
  • Coordinate with cross-functional teams during incident resolution
  • Participate in major incident bridges, DR exercises, and 24x7 operations support activities
  • Follow escalation procedures, SOPs, operational runbooks, and ITIL processes

Instana & APM Operations:

  • Administer and support IBM Instana monitoring environments
  • Monitor application, API, middleware, and microservices performance using Instana
  • Validate Instana agent health following server patching and maintenance activities
  • Configure alerts, baselines, and performance thresholds

Automation & Scripting:

  • Develop automation solutions using Python, PowerShell, Shell Scripting (Bash), and VBScript
  • Automate operational tasks, monitoring deployments, and remediation workflows
  • Build reusable automation tools to improve operational efficiency
  • Integrate monitoring platforms with enterprise automation frameworks

Configuration Management & Infrastructure Automation:

  • Implement Infrastructure as Code (IaC) and automation using Ansible and Puppet
  • Automate server provisioning and configuration management
  • Automate monitoring agent deployment and onboarding
  • Maintain automation playbooks and deployment pipelines

Knowledge Management:

  • Maintain SOPs, runbooks, monitoring procedures, and escalation matrices
  • Participate in knowledge transfer sessions, service onboarding, and operational readiness reviews
  • Support service transition, migration, and continuous improvement initiatives

What We're Looking For

Required Skills:

  • 7 to 10+ years of experience in Monitoring, Observability, Infrastructure Operations, SRE, or Platform Engineering
  • Hands-on experience with Grafana, IBM Instana, SolarWinds, Telegraf, Prometheus, and InfluxDB
  • Experience supporting large-scale enterprise environments and 24x7 operations
  • Expertise in VMware, Linux Administration, Windows Server, Citrix VDI, DNS Services, Proxy Services, and Middleware Technologies
  • Strong scripting skills in Python, PowerShell, Shell Scripting (Linux), and VBScript
  • Experience with automation tools including Ansible and Puppet
  • FOAK application support and Enterprise Logging & Telemetry (ELT) experience
  • Experience with ServiceNow, Incident Management, Problem Management, Change Management, and ITIL Framework
  • Strong troubleshooting, RCA, and operational support skills
  • Experience integrating observability platforms with enterprise automation solutions

Preferred Skills:

  • Docker and Kubernetes
  • Cloud platforms: AWS, Microsoft Azure, Google Cloud Platform (GCP)
  • CI/CD tools: Jenkins, GitHub Actions, GitLab CI/CD
  • REST APIs and Microservices Monitoring
  • DevOps & SRE Practices

How We Work With You

When you apply through Get A Job.ai, our recruiting team will review your qualifications and experience against our client's requirements. If we determine there's a strong match, we'll conduct an initial screening conversation to discuss your background, career goals, and the role in detail.

After our screening, we will submit qualified candidates to our client for consideration. Throughout the interview process, we remain your advocate and main point of contact. Please do not contact the client organization directly—all communication should flow through Get A Job.ai to ensure a smooth and professional process.

Pay

Pay details will be discussed during the screening process based on your experience and qualifications.

Equal Employment Opportunity: Get A Job.ai is committed to providing equal employment opportunities to all applicants. We evaluate candidates based on qualifications and merit without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, veteran status, or other legally protected characteristics.

Apply with Get A Job.ai

A recruiter will review your profile and submit you to the client. Do not contact the client directly.

More options

Apply with Get A Job.ai

Apply through Get A Job.ai. A recruiter will review your profile and submit you. Do not contact the client directly.

Apply through Get A Job.ai. A recruiter will review your profile and submit you.

Working in United States

The United States of America (USA), also known as the United States (U.S.) or America, is a country primarily located in North America. It is a federal republic consisting of 50 states and a federal capital district, Washington, D.C. The 48 contiguous states border Canada to the north and Mexico to the south, with the semi-exclave of Alaska in the northwest and the archipelago of Hawaii in the Pacific Ocean. The United States also asserts sovereignty over five major island territories and various uninhabited islands in Oceania and the Caribbean. It is a megadiverse country, with the world's th

🇺🇸 Relocation safety for US: Exercise Normal Cautionvia Warnely, CC BY 4.0

National unemployment rate in US: 4.2%via World Bank

National job openings rate: 4.4%via BLS JOLTS

Private-sector wage growth (year over year): 3.1%via FRED

National quits rate: 2.0%via FRED (BLS JOLTS)

Weekly initial unemployment claims: 206,000via FRED

GDP per capita in US: $90,027via World Bank

Consumer price inflation in US: 2.9% (annual) — via World Bank

Real GDP growth in US: 2.2% (annual) — via World Bank

Average hours worked per year in US: 1,800via OECD

    Job details above are provided by the employer/source. The sections on this page are compiled from public data sources with AI assistance.

    Accommodations: if you need a workplace accommodation to apply for or perform this job, see ADA.gov or EEOC.gov for guidance on your rights and how to request one.

    Add application deadline to calendar

    Listing facts

    • Role Senior Automation & Observability Engineer
    • Employer Get A Job.ai
    • Location United States
    • Type Full Time
    • Pay (from listing) Pay not listed
    • Posted September 19, 2026
    • Apply by October 19, 2026
    • Country United States
    • Overview Full job description on this page (779 words)

    Facts above come from this job record on Get A Job.AI — not copied from third-party review sites.

    Limited public data for this employer

    We only show facts we can ground in public sources (Wikidata, O*NET, news/discussion links, or this listing). We do not invent Glassdoor-style ratings, salaries, or testimonials when data is thin. Use the listing facts, occupation context, and related openings below while we continue researching.

    Explore related openings

    Keep exploring on Get A Job.ai

    Not quite the right fit? Your next opportunity is a click away.

    Hiring instead? Post a job and reach candidates searching right now.

    Senior Automation & Observability Engineer Get A Job.ai · United States