- Company: Get A Job.ai
- Location: United States
- Salary: Pay not listed
Website Get A Job.ai
Represented by Get A Job.ai
About This Opportunity
Our talent team at Get A Job.ai is representing a confidential online gaming organization seeking a Senior Machine Learning Operations Engineer to join their platform engineering group in the United States. This is a remote position where you'll own the full production lifecycle of ML systems—from trained models to production endpoints—treating machine learning infrastructure as software systems with reliability, cost, and latency requirements.
You'll build the platform that data scientists and ML engineers ship on: feature stores with guaranteed online/offline parity, model registries, CI/CD pipelines for ML, drift monitoring, and deployment scaffolding for champion/challenger patterns. This role demands a software-engineering-first mindset with distributed systems expertise and on-call experience; ML literacy makes you effective in this domain.
Responsibilities
ML Production Platform
- Stand up and operate the ML platform on AWS (SageMaker Training, Model Registry, Pipelines, Endpoints, Batch Transform) and Snowflake (Snowpark ML, Cortex) with Terraform-managed infrastructure
- Build self-service scaffolds enabling data scientists to ship models end-to-end without ticket queues—cookie-cutter templates with CI, drift monitoring, alerting, IaC, and Snowflake connectivity pre-configured
Batch and Real-Time Inference
- Design and operate batch scoring pipelines—SageMaker Batch Transform, dbt-orchestrated scoring against Snowflake, Snowpark ML—with explicit freshness and cost SLAs
- Design and operate real-time inference paths—SageMaker endpoints, Lambda + Bedrock for GenAI, API Gateway—with stated latency budgets (typically sub-100ms) and graceful degradation under load
- Own the feature store (SageMaker Feature Store, Tecton, or Feast) with guaranteed online/offline parity; training-serving skew is treated as an incident
CI/CD and Deployment Patterns
- Build CI/CD for ML: model registry, automated retraining triggers, model versioning, lineage from feature → training run → deployed model → live prediction
- Implement champion/challenger, shadow deployments, and canary releases as platform primitives
Monitoring, Drift & Reliability
- Stand up drift detection, data quality, and model performance monitoring (Evidently, Arize, or SageMaker Model Monitor) with paging that routes to engineers who can respond
- Own MLOps incident response; production model failures are SEV events with postmortems
Cost and Performance
- Right-size endpoints, implement batch caching, request batching, and autoscaling; state cost-per-prediction targets and meet them
GenAI Integration (Plus, Not Required)
- Integrate LLM APIs (Bedrock, Anthropic, OpenAI) into production paths—RAG pipelines, agent eval frameworks, prompt versioning, cost and latency observability
- Direct AI coding agents (Claude Code, Cursor, GitHub Copilot, dbt Copilot) as force multipliers across infrastructure code and model-serving glue
Collaboration
- Partner with data engineering teams on shared standards (Terraform modules, CI/CD patterns, observability, lineage)
- Work alongside data scientists and analytics partners to define the right interfaces between research and production
What We're Looking For
Education: BS or MS in Computer Science, Math, Statistics, Machine Learning, or other STEM field—or equivalent practical experience
Must-Have Experience:
- 5+ years shipping software in production—Python, Docker, Kubernetes or ECS, CI/CD, distributed systems debugging—including on-call time
- 3+ years operating ML in production—you have owned a model serving real traffic with stated latency and cost budgets and a runbook you wrote
- AWS depth across SageMaker (Training, Endpoints, Batch Transform, Model Registry, Pipelines) plus IAM, Lambda, ECS, S3, Secrets Manager, VPC
- Snowflake fluency—Snowpark ML, Cortex, dbt-orchestrated batch scoring, RBAC for ML workloads
- IaC for ML—Terraform + SageMaker Pipelines or equivalent; no manual console deployments to production
- Feature store experience—SageMaker Feature Store, Tecton, or Feast—with explicit ownership of online/offline parity
- Champion/challenger, shadow, and canary deployment patterns as production muscle
- Drift and model monitoring—Evidently, Arize, WhyLabs, or SageMaker Model Monitor—wired to a paging path
- Software-engineering-first mindset—you treat ML systems as systems, not notebooks
Nice-to-Have Experience:
- GenAI in production—Bedrock, Anthropic, or OpenAI APIs integrated into live systems; RAG pipelines; vector DBs; evaluation frameworks
- Snowflake-native ML—Snowpark Container Services, Cortex AI SQL, Cortex Agents
- Streaming feature engineering—Kafka, Flink, or Snowpipe Streaming—for sub-second features
- Fine-tuning experience—LoRA, QLoRA, instruction tuning, eval-driven iteration
- A track record of shipping more with AI in the engineering loop than without
- Regulated-industry experience (gaming, fintech, healthcare)—comfort with model risk, audit, and lineage requirements
Legal Requirements: Candidates must possess legal authorization to work in the U.S. without immigration sponsorship. This role is not eligible for H-1B, O-1, E-3, TN, OPT, or other immigration-related work authorization. Gaming compliance and licensing requirements apply, including comprehensive background checks covering criminal records, financial history, and personal background verification.
How We Work With You
When you apply through Get A Job.ai, our recruiting team will review your profile and schedule an initial screening call to understand your background and career goals. If there's a strong match, we'll submit your candidacy to our client for consideration and guide you through their interview process. Please do not contact the client directly—all communication should flow through Get A Job.ai to ensure a coordinated and professional experience.
Pay
The client has budgeted an annual salary range of $135,000 to $170,000 for this position. Factors affecting starting pay within this range include geography, skills, education, experience, and other qualifications. This position is also eligible for participation in a performance-based bonus plan.
Get A Job.ai is an equal opportunity recruiter. We welcome candidates from all backgrounds and provide equal consideration regardless of race, religion, gender, gender identity, age, marital status, national origin, sexual orientation, citizenship status, veteran status, disability, or any other legally protected status.
Apply with Get A Job.ai
A recruiter will review your profile and submit you to the client. Do not contact the client directly.
Apply through Get A Job.ai. A recruiter will review your profile and submit you. Do not contact the client directly.
Apply through Get A Job.ai. A recruiter will review your profile and submit you.
Local insights for this role are preparing — this section updates automatically in a few seconds (or refresh).
Listing facts
- Role Senior Machine Learning Operations Engineer
- Employer Get A Job.ai
- Location United States
- Type Full Time
- Pay (from listing) Pay not listed
- Posted September 11, 2026
- Apply by October 12, 2026
- Country United States
- Overview Full job description on this page (879 words)
Facts above come from this job record on Get A Job.AI — not copied from third-party review sites.
Limited public data for this employer
We only show facts we can ground in public sources (Wikidata, O*NET, news/discussion links, or this listing). We do not invent Glassdoor-style ratings, salaries, or testimonials when data is thin. Use the listing facts, occupation context, and related openings below while we continue researching.
Explore related openings
Keep exploring on Get A Job.ai
Not quite the right fit? Your next opportunity is a click away.
- Browse all jobs
- More jobs by category
- Remote jobs you can do from anywhere
- Research typical pay for this role
- Set a job alert so new matches reach you first
- Upload your resume to apply faster
Hiring instead? Post a job and reach candidates searching right now.
