- Company: Get A Job.ai
- Location: Canada
- Salary: Pay not listed
- Work type: Remote
Website Get A Job.ai
Represented by Get A Job.ai
About This Opportunity
We are representing a confidential client in the customer engagement technology sector, seeking a Senior Site Reliability Engineer to join their Canadian operations. This is a critical infrastructure role focused on maintaining and scaling the ingress layers that serve as the primary entry point for massive API traffic volumes.
Our client operates distributed systems at extraordinary scale, and this position will be embedded with product engineering teams to ensure their Ruby on Rails and Go-based API services remain highly available and resilient under extreme load conditions.
Responsibilities
- Lead NGINX & Kubernetes Ingress Infrastructure: Architect, configure, and operate high-performance NGINX routing, proxying, and ingress controller layers managing real-time API ingestion traffic at scale
- Own Scaling Operations: Design, expand, and tune automated scaling routines for high-throughput API services using RED metrics, Horizontal Pod Autoscalers, and custom scaling policies to handle traffic spikes seamlessly
- Partner with Product Teams: Translate product requirements into resilient, highly available technology stacks, working directly with engineering teams owning API features
- Establish SLIs and SLOs: Define meaningful Service Level Indicators and Objectives for API services, helping teams balance rapid feature deployment with stability through error budget management
- Systems Design & Capacity Planning: Conduct bottleneck profiling and capacity planning to ensure enterprise-grade SLA compliance
- Incident Response: Participate in PagerDuty on-call rotation, lead root-cause analysis and blameless retrospectives, and translate operational learnings into permanent system improvements
- Proactive Prevention: Update runbooks and implement preventive measures to reduce recurring incidents
What We're Looking For
Required Experience & Skills:
- 5+ years of experience as a Site Reliability Engineer or DevOps Engineer in high-scale production environments
- Deep hands-on expertise configuring, troubleshooting, and operating NGINX proxying, routing, and ingress controllers under heavy traffic loads
- In-depth proficiency with Kubernetes administration, including cluster networking, container orchestration, scheduling, and deployment
- Excellent understanding of Linux/Unix internals: disk I/O, memory allocation, TCP/IP networking, and process management
- Strong programming/scripting skills in Ruby and/or Go preferred (Python or Java also valued) for building automation tools and platform frameworks
- Experience managing infrastructure using IaC technologies such as Terraform, Ansible, or Chef
- Strong systems thinking: understanding interfaces, boundaries, failure modes, and cascading effects in distributed architectures
- Excellent documentation practices and comfort collaborating asynchronously across remote-first, global teams
Preferred Skills:
- Familiarity with data-tier technologies including Redis, Kafka, Postgres, or MongoDB
- Experience with observability platforms such as Prometheus, Grafana, or Datadog
- Practical experience deploying infrastructure in major cloud environments (AWS, GCP, Azure)
How We Work With You
Our talent team at Get A Job.ai specializes in connecting top infrastructure engineers with confidential opportunities at leading technology companies. Here's our process:
- Apply directly through the Get A Job.ai platform
- A recruiter from our team will conduct an initial screening to understand your background and goals
- We will prepare and submit your profile to our client for consideration
- Please do not attempt to contact the client directly, as all communication is managed through Get A Job.ai to ensure a smooth process for all parties
Pay
This role is based in Canada. Compensation details will be discussed during the screening process with our recruiting team.
Equal Opportunity: Get A Job.ai is committed to fostering an inclusive recruitment process. We welcome applications from candidates of all backgrounds and experiences, and we're happy to discuss any accommodations you may need during the interview process.
Apply with Get A Job.ai
A recruiter will review your profile and submit you to the client. Do not contact the client directly.
Apply through Get A Job.ai. A recruiter will review your profile and submit you. Do not contact the client directly.
Apply through Get A Job.ai. A recruiter will review your profile and submit you.
Local insights for this role are preparing — this section updates automatically in a few seconds (or refresh).
Listing facts
- Role Senior Site Reliability Engineer I
- Employer Get A Job.ai
- Location Canada · Remote-friendly
- Type Full Time
- Pay (from listing) Pay not listed
- Posted September 4, 2026
- Apply by October 4, 2026
- Country Canada
- Overview Full job description on this page (553 words)
Facts above come from this job record on Get A Job.AI — not copied from third-party review sites.
Typical work in Senior Site Reliability Engineer
Independent occupational context from O*NET (U.S. public-domain labor data). This is about the occupation, not a rewrite of this employer's posting.
- Study product characteristics or customer requirements to determine validation objectives and standards.
- Analyze validation test data to determine whether systems or processes have met validation criteria or to identify root causes of production problems.
- Develop validation master plans, process flow diagrams, test cases, or standard operating procedures.
- Prepare detailed reports or design statements, based on results of validation and qualification tests or reviews of procedures and protocols.
- Maintain validation test equipment.
- Conduct validation or qualification tests of new or existing processes, equipment, or software in accordance with internal protocols or external standards.
Limited public data for this employer
We only show facts we can ground in public sources (Wikidata, O*NET, news/discussion links, or this listing). We do not invent Glassdoor-style ratings, salaries, or testimonials when data is thin. Use the listing facts, occupation context, and related openings below while we continue researching.
Explore related openings
Occupation family: Senior Site Reliability Engineer
Keep exploring on Get A Job.ai
Not quite the right fit? Your next opportunity is a click away.
- Browse all jobs
- More jobs by category
- Remote jobs you can do from anywhere
- Research typical pay for this role
- Set a job alert so new matches reach you first
- Upload your resume to apply faster
Hiring instead? Post a job and reach candidates searching right now.
