Loading...

Senior Machine Learning Operations Engineer

  • Company: Get A Job.ai
  • Location: United States
  • Salary: Pay not listed

Website Get A Job.ai

Represented by Get A Job.ai

About This Opportunity

Our talent team at Get A Job.ai is representing a confidential online gaming organization seeking a Senior Machine Learning Operations Engineer to join their platform engineering group in the United States. This is a remote position where you'll own the full production lifecycle of ML systems—from trained models to production endpoints—treating machine learning infrastructure as software systems with reliability, cost, and latency requirements.

You'll build the platform that data scientists and ML engineers ship on: feature stores with guaranteed online/offline parity, model registries, CI/CD pipelines for ML, drift monitoring, and deployment scaffolding for champion/challenger patterns. This role demands a software-engineering-first mindset with distributed systems expertise and on-call experience; ML literacy makes you effective in this domain.

Responsibilities

ML Production Platform

  • Stand up and operate the ML platform on AWS (SageMaker Training, Model Registry, Pipelines, Endpoints, Batch Transform) and Snowflake (Snowpark ML, Cortex) with Terraform-managed infrastructure
  • Build self-service scaffolds enabling data scientists to ship models end-to-end without ticket queues—cookie-cutter templates with CI, drift monitoring, alerting, IaC, and Snowflake connectivity pre-configured

Batch and Real-Time Inference

  • Design and operate batch scoring pipelines—SageMaker Batch Transform, dbt-orchestrated scoring against Snowflake, Snowpark ML—with explicit freshness and cost SLAs
  • Design and operate real-time inference paths—SageMaker endpoints, Lambda + Bedrock for GenAI, API Gateway—with stated latency budgets (typically sub-100ms) and graceful degradation under load
  • Own the feature store (SageMaker Feature Store, Tecton, or Feast) with guaranteed online/offline parity; training-serving skew is treated as an incident

CI/CD and Deployment Patterns

  • Build CI/CD for ML: model registry, automated retraining triggers, model versioning, lineage from feature → training run → deployed model → live prediction
  • Implement champion/challenger, shadow deployments, and canary releases as platform primitives

Monitoring, Drift & Reliability

  • Stand up drift detection, data quality, and model performance monitoring (Evidently, Arize, or SageMaker Model Monitor) with paging that routes to engineers who can respond
  • Own MLOps incident response; production model failures are SEV events with postmortems

Cost and Performance

  • Right-size endpoints, implement batch caching, request batching, and autoscaling; state cost-per-prediction targets and meet them

GenAI Integration (Plus, Not Required)

  • Integrate LLM APIs (Bedrock, Anthropic, OpenAI) into production paths—RAG pipelines, agent eval frameworks, prompt versioning, cost and latency observability
  • Direct AI coding agents (Claude Code, Cursor, GitHub Copilot, dbt Copilot) as force multipliers across infrastructure code and model-serving glue

Collaboration

  • Partner with data engineering teams on shared standards (Terraform modules, CI/CD patterns, observability, lineage)
  • Work alongside data scientists and analytics partners to define the right interfaces between research and production

What We're Looking For

Education: BS or MS in Computer Science, Math, Statistics, Machine Learning, or other STEM field—or equivalent practical experience

Must-Have Experience:

  • 5+ years shipping software in production—Python, Docker, Kubernetes or ECS, CI/CD, distributed systems debugging—including on-call time
  • 3+ years operating ML in production—you have owned a model serving real traffic with stated latency and cost budgets and a runbook you wrote
  • AWS depth across SageMaker (Training, Endpoints, Batch Transform, Model Registry, Pipelines) plus IAM, Lambda, ECS, S3, Secrets Manager, VPC
  • Snowflake fluency—Snowpark ML, Cortex, dbt-orchestrated batch scoring, RBAC for ML workloads
  • IaC for ML—Terraform + SageMaker Pipelines or equivalent; no manual console deployments to production
  • Feature store experience—SageMaker Feature Store, Tecton, or Feast—with explicit ownership of online/offline parity
  • Champion/challenger, shadow, and canary deployment patterns as production muscle
  • Drift and model monitoring—Evidently, Arize, WhyLabs, or SageMaker Model Monitor—wired to a paging path
  • Software-engineering-first mindset—you treat ML systems as systems, not notebooks

Nice-to-Have Experience:

  • GenAI in production—Bedrock, Anthropic, or OpenAI APIs integrated into live systems; RAG pipelines; vector DBs; evaluation frameworks
  • Snowflake-native ML—Snowpark Container Services, Cortex AI SQL, Cortex Agents
  • Streaming feature engineering—Kafka, Flink, or Snowpipe Streaming—for sub-second features
  • Fine-tuning experience—LoRA, QLoRA, instruction tuning, eval-driven iteration
  • A track record of shipping more with AI in the engineering loop than without
  • Regulated-industry experience (gaming, fintech, healthcare)—comfort with model risk, audit, and lineage requirements

Legal Requirements: Candidates must possess legal authorization to work in the U.S. without immigration sponsorship. This role is not eligible for H-1B, O-1, E-3, TN, OPT, or other immigration-related work authorization. Gaming compliance and licensing requirements apply, including comprehensive background checks covering criminal records, financial history, and personal background verification.

How We Work With You

When you apply through Get A Job.ai, our recruiting team will review your profile and schedule an initial screening call to understand your background and career goals. If there's a strong match, we'll submit your candidacy to our client for consideration and guide you through their interview process. Please do not contact the client directly—all communication should flow through Get A Job.ai to ensure a coordinated and professional experience.

Pay

The client has budgeted an annual salary range of $135,000 to $170,000 for this position. Factors affecting starting pay within this range include geography, skills, education, experience, and other qualifications. This position is also eligible for participation in a performance-based bonus plan.

Get A Job.ai is an equal opportunity recruiter. We welcome candidates from all backgrounds and provide equal consideration regardless of race, religion, gender, gender identity, age, marital status, national origin, sexual orientation, citizenship status, veteran status, disability, or any other legally protected status.

Apply with Get A Job.ai

A recruiter will review your profile and submit you to the client. Do not contact the client directly.

More options

Apply with Get A Job.ai

Apply through Get A Job.ai. A recruiter will review your profile and submit you. Do not contact the client directly.

Apply through Get A Job.ai. A recruiter will review your profile and submit you.

Terms used in this posting

on-call
You may be required to be reachable and available to work outside normal scheduled hours, typically for a set rotation.

Working in United States

The United States of America (USA), also known as the United States (U.S.) or America, is a country primarily located in North America. It is a federal republic consisting of 50 states and a federal capital district, Washington, D.C. The 48 contiguous states border Canada to the north and Mexico to the south, with the semi-exclave of Alaska in the northwest and the archipelago of Hawaii in the Pacific Ocean. The United States also asserts sovereignty over five major island territories and various uninhabited islands in Oceania and the Caribbean. It is a megadiverse country, with the world's th

🇺🇸 Relocation safety for US: Exercise Normal Cautionvia Warnely, CC BY 4.0

National unemployment rate in US: 4.2%via World Bank

National job openings rate: 4.4%via BLS JOLTS

Private-sector wage growth (year over year): 3.1%via FRED

National quits rate: 2.0%via FRED (BLS JOLTS)

Weekly initial unemployment claims: 196,000via FRED

GDP per capita in US: $90,027via World Bank

Consumer price inflation in US: 2.9% (annual) — via World Bank

Real GDP growth in US: 2.2% (annual) — via World Bank

Average hours worked per year in US: 1,800via OECD

    Job details above are provided by the employer/source. The sections on this page are compiled from public data sources with AI assistance.

    Accommodations: if you need a workplace accommodation to apply for or perform this job, see ADA.gov or EEOC.gov for guidance on your rights and how to request one.

    Add application deadline to calendar

    Listing facts

    • Role Senior Machine Learning Operations Engineer
    • Employer Get A Job.ai
    • Location United States
    • Type Full Time
    • Pay (from listing) Pay not listed
    • Posted September 11, 2026
    • Apply by October 12, 2026
    • Country United States
    • Overview Full job description on this page (879 words)

    Facts above come from this job record on Get A Job.AI — not copied from third-party review sites.

    Limited public data for this employer

    We only show facts we can ground in public sources (Wikidata, O*NET, news/discussion links, or this listing). We do not invent Glassdoor-style ratings, salaries, or testimonials when data is thin. Use the listing facts, occupation context, and related openings below while we continue researching.

    Explore related openings

    Keep exploring on Get A Job.ai

    Not quite the right fit? Your next opportunity is a click away.

    Hiring instead? Post a job and reach candidates searching right now.

    Senior Machine Learning Operations Engineer Get A Job.ai · United States