Loading...

Principal Infrastructure Engineer

  • Company: Get A Job.ai
  • Salary: Pay not listed
  • Work type: Remote

Website Get A Job.ai

Represented by Get A Job.ai

Responsibilities

We are representing a confidential fintech organization seeking a Principal Infrastructure Engineer to design, build, and scale the platform that powers their payment operations. You will tackle the hardest infrastructure challenges: increasing throughput, reducing latency, removing capacity bottlenecks, strengthening resilience, and making production operations more automated and predictable as the business grows.

Your responsibilities include:

  • Own the technical architecture and evolution of core infrastructure: identify system limits, prioritize technical improvements, and implement changes that support increasing traffic, data volume, and workload complexity.
  • Connect business understanding to improvements across the system: learn how key business workflows behave in production, trace their impact across applications, data, and infrastructure, and partner across teams to improve performance, reliability, and cost efficiency.
  • Engineer for scale and performance: build capacity models, run load and stress tests, diagnose bottlenecks across compute, networking, Kubernetes, and databases, and validate improvements against throughput, latency, saturation, and cost per workload.
  • Design and build AWS infrastructure: implement resilient account, IAM, network, and service architectures; address service quotas, fault isolation, and multi-AZ or multi-region requirements as workloads grow.
  • Build and operate the Kubernetes platform: improve cluster architecture, lifecycle automation, workload isolation, resource allocation, autoscaling, safe upgrades, and deployment reliability.
  • Scale and optimize Aurora RDS for MySQL and Postgres: tune queries and indexes, address connection and replication bottlenecks, plan capacity, improve failover behavior, and implement safe schema changes and database migrations with application engineers.
  • Improve reliability through engineering: define and instrument service-level objectives and error budgets with service owners; implement failure isolation, backpressure, load shedding, and safe retry behavior where needed to prevent cascading failures.
  • Participate in on-call and drive technical recovery during serious incidents: use evidence-based triage, execute mitigations, communicate findings, and implement corrective actions from postmortems.
  • Implement and test disaster recovery: design backup, restore, and failover mechanisms against agreed recovery time and recovery point objectives; run recovery exercises and document measured results.
  • Build infrastructure-as-code and operational automation: make provisioning, configuration, deployments, upgrades, and recovery reproducible, reviewed, testable, and recoverable.
  • Build observability that makes production diagnosable: improve metrics, logs, traces, dashboards, and actionable alerts, with visibility into service health, scaling limits, and customer impact.
  • Deliver safe infrastructure migrations: design phased rollouts, compatibility checks, validation, and rollback paths for changes to shared production systems.
  • Improve cloud cost efficiency through technical changes: right-size resources, improve utilization, tune autoscaling and storage, and quantify savings while maintaining reliability and performance targets.
  • Build and evaluate AI-assisted operational tooling: apply AI to investigation, runbooks, anomaly analysis, and repetitive operations, with bounded permissions, reviewable actions, and measurable improvements in accuracy or toil.

What We're Looking For

Required qualifications:

  • Bachelor's degree in Computer Science or a similar technical field
  • 12+ years of experience across infrastructure, platform, site reliability, software development, or related engineering disciplines, with substantial hands-on depth designing and operating production infrastructure at scale
  • Deep expertise with AWS: production experience across compute, IAM, multi-account architectures, and networking, including VPC design and private connectivity
  • Deep expertise with Kubernetes in production: cluster lifecycle, scheduling, resource management, autoscaling, networking, and troubleshooting business-critical workloads. EKS experience is strongly preferred
  • Deep expertise with relational databases at scale, specifically RDS/Aurora (MySQL and/or Postgres): query performance, indexing, connection management, replication, high availability, failover, and verified backup and recovery
  • A track record of personally delivering infrastructure scaling improvements: identifying constraints, measuring baseline behavior, implementing changes, and demonstrating gains in capacity, latency, reliability, or cost efficiency
  • Strong coding and automation skills, using Golang, Python, or similar languages to build production tooling and eliminate operational toil, alongside infrastructure-as-code experience with Terraform or equivalent
  • Strong systems fundamentals: Linux, networking, DNS, TLS, storage, concurrency, and distributed system failure modes, with the ability to debug problems across infrastructure and application boundaries
  • Experience operating a 24/7, high-availability platform where downtime has direct customer or revenue impact, including hands-on incident response and postmortem remediation
  • Willingness to participate in an on-call rotation and demonstrated ability to recover production systems under pressure using evidence-based triage, safe mitigation, and clear technical communication
  • Experience implementing and testing disaster recovery against defined recovery objectives, including restoring data and validating service recovery
  • Practical experience with observability, load testing, capacity planning, and safe CI/CD practices for shared production infrastructure
  • Active use of AI tooling in engineering or operations, with practical judgment about its limitations and how to verify generated code, recommendations, and operational actions
  • Ability to carry ambiguous technical problems from investigation through production delivery and collaborate across engineering disciplines to resolve system-wide constraints

Preferred qualifications:

  • Experience in fintech, payments, or banking, operating infrastructure with demanding reliability, security, and audit requirements
  • Experience with multi-region architectures, chaos engineering, and failure testing, including the consistency and recovery tradeoffs of distributed data systems
  • Proficiency with Prometheus, Grafana, Loki, Tempo, or comparable observability systems
  • Experience building internal platform capabilities and self-service tooling, including deployment automation, progressive delivery, and reusable infrastructure components
  • Experience building AI-assisted incident investigation or operational automation with restricted access, auditable execution, and clear human review points

How We Work With You

Get A Job.ai represents exceptional technology organizations seeking senior talent. When you apply through our platform, our talent team will conduct an initial screening to understand your background and career goals. If there's a strong match, we will submit your profile to our client for consideration. Throughout the process, we serve as your advocate and guide.

Please apply exclusively through Get A Job.ai. Do not attempt to contact the client directly, as this may disqualify your application.

Pay

This is a remote, full-time position. The technology stack includes Golang, Python, MySQL, Postgres, AWS, Kubernetes, Git, and GitLab. Our client values open-source solutions and builds custom tooling where appropriate.

Get A Job.ai is an equal opportunity recruiting firm. We welcome applications from candidates of all backgrounds and are committed to fostering an inclusive hiring process.

Apply with Get A Job.ai

A recruiter will review your profile and submit you to the client. Do not contact the client directly.

More options

Apply with Get A Job.ai

Apply through Get A Job.ai. A recruiter will review your profile and submit you. Do not contact the client directly.

Apply through Get A Job.ai. A recruiter will review your profile and submit you.

Terms used in this posting

on-call
You may be required to be reachable and available to work outside normal scheduled hours, typically for a set rotation.

About this role & career path

Traits that fit this role

  • Adaptability
  • Innovation
  • Intellectual Curiosity
  • Cautiousness
  • Attention to Detail

Source: O*NET Work Styles (Distinctiveness Rank).

Typical preparation needed: Job Zone 3: Medium Preparation Needed. Most occupations in this zone require training in vocational schools, related on-the-job experience, or an associate's degree. — via O*NET

Industry news

Source: O*NET (public-domain bulk data)

Salary & compensation

Workers in Computer & Mathematical occupations earn a national median of $98,769via US Census ACS / Data USA

Market context

  • Similar listings in Remote 1

Job details above are provided by the employer/source. The sections on this page are compiled from public data sources with AI assistance.

Accommodations: if you need a workplace accommodation to apply for or perform this job, see ADA.gov or EEOC.gov for guidance on your rights and how to request one.

Add application deadline to calendar

Listing facts

  • Role Principal Infrastructure Engineer
  • Employer Get A Job.ai
  • Location Remote · Remote-friendly
  • Type Full Time
  • Pay (from listing) Pay not listed
  • Posted September 20, 2026
  • Apply by October 20, 2026
  • Overview Full job description on this page (957 words)

Facts above come from this job record on Get A Job.AI — not copied from third-party review sites.

Typical work in Infrastructure Engineer

Independent occupational context from O*NET (U.S. public-domain labor data). This is about the occupation, not a rewrite of this employer's posting.

  • Consult with users, administrators, and engineers to identify business and technical requirements for proposed system modifications or technology purchases.
  • Implement system renovation projects in collaboration with technical staff, engineering consultants, installers, and vendors.
  • Keep abreast of changes in industry practices and emerging telecommunications technology by reviewing current literature, talking with colleagues, participating in educational programs, attending meetings or workshops, or participating in professional organizations or conferences.
  • Review and evaluate requests from engineers, managers, and technicians for system modifications.
  • Assess existing facilities' needs for new or modified telecommunications systems.
  • Develop, maintain, or implement telecommunications disaster recovery plans to ensure business continuity.

Source: O*NET

Limited public data for this employer

We only show facts we can ground in public sources (Wikidata, O*NET, news/discussion links, or this listing). We do not invent Glassdoor-style ratings, salaries, or testimonials when data is thin. Use the listing facts, occupation context, and related openings below while we continue researching.

Explore related openings

Occupation family: Infrastructure Engineer

Keep exploring on Get A Job.ai

Not quite the right fit? Your next opportunity is a click away.

Hiring instead? Post a job and reach candidates searching right now.

Principal Infrastructure Engineer Get A Job.ai · Remote