Print / Save PDF Back to job

Site Reliability Engineer – Azure, Observability and Scripting

Jalasoft · Colombia

How to use this kit

Ground every answer in facts on this page and the original listing. We never invent Glassdoor-style reviews or salaries that are not in our data.

Interview prep

Expect deep dives on Azure services, monitoring stacks, and scripting (PowerShell, Bash, or Python). Be ready to walk through an incident: how you triaged signals, mitigated impact, and prevented recurrence. Prepare examples of automation that cut MTTR or manual ops, plus trade-offs between alerting noise and coverage.

Fit summary

Strong fit if you enjoy production ownership, Azure operations, and turning repeat fixes into scripts and observability. Less ideal if you prefer pure feature development with no on-call or reliability metrics.

Day in the role

As a Site Reliability Engineer focused on Azure, observability, and scripting at Jalasoft, a typical day centers on keeping cloud services healthy: reviewing alerts and dashboards, tuning monitors and SLOs, and automating toil with scripts. You would diagnose incidents, harden deployments on Azure, and partner with product teams so reliability work ships with features rather than after outages.

FAQ from this listing

Is this role remote?

The listing marks remote as no and places the job in Colombia, so plan for local or employer-defined on-site/hybrid arrangements rather than fully remote work.

What skills does the title emphasize?

Azure platform reliability, observability (metrics, logs, traces, alerting), and scripting for automation and incident response.

What would success look like early on?

Fewer noisy alerts, clearer runbooks, automated routine checks, and faster recovery when Azure-hosted services degrade.

Role overview (listing rewrite)

SITE RELIABILITY ENGINEER - AZURE, OBSERVABILITY AND SCRIPTINGAt Jalasoft in ColombiaFull Time, No Remote Position Summary We are seeking a Site Reliability Engineer (SRE) to join our team and ensure the reliability, scalability, and performance of cloud-native platforms running on Microsoft Azure and Kubernetes. This role requires passion for observability, automation, and operational excellence. Day-to-Day Responsibilities Improve system availability, monitoring, incident response, and infrastructure automation in production environments Operate Kubernetes in production at scale Utilize Azure Monitor, Log Analytics, KQL, Prometheus, Grafana, and other tools for metrics and alerting Define and implement service level indicators, objectives, and error budgets Design and test backup, restore, and disaster recovery strategies Troubleshoot Kubernetes workloads, manage resources, and upgrade clusters Read and modify infrastructure as code (Terraform or Bicep) and Azure DevOps pipelines Create and maintain scripts in Python, PowerShell, or Bash Skills & Qualifications 6+ years of experience in Site Reliability Engineering or Production Operations for Kubernetes workloads at scale 3+ years of experience with Azure Monitor, Log Analytics, and KQL Experience with Prometheus, Grafana, and other monitoring tools Proficiency in defining and implementing service level indicators, objectives, and error budgets Knowledge of alerting and incident response design, including runbook authoring and on-call practice Experience with backup, restore, and disaster recovery design and testing Familiarity with Kubernetes operations: workload troubleshooting, resource management, and cluster upgrades Able to read and modify infrastructure as code (Terraform or Bicep) and Azure DevOps pipelines Scripting skills in Python, PowerShell, or Bash Professional working English Nice-to-have: Azure Managed Prometheus and Grafana, OpenTelemetry instrumentation, database-layer observability for Oracle, cost and capacity management for AKS estates, incident management tooling More About Jalasoft Jalasoft is a leading technology company dedicated to delivering innovative solutions that empower businesses. Our team of experts works collaboratively to solve complex challenges in the tech industry. Apply Today To apply, complete your application directly on this page, or you'll be redirected to the employer's application platform to finish submitting there.

Full job on Get A Job.AI

Questions to ask them

Generated for personal interview prep · 2026-08-09 UTC · getajob.ai