- Company: Get A Job.ai
- Salary: Pay not listed
- Work type: Remote
Website Get A Job.ai
Represented by Get A Job.ai
Responsibilities
We are representing a confidential fintech organization seeking a Principal Infrastructure Engineer to design, build, operate, and scale their cloud platform. You will tackle the hardest infrastructure challenges: increasing throughput, reducing latency, removing capacity bottlenecks, strengthening resilience, and making production operations more automated and predictable as the business grows.
You will own complex technical initiatives from architecture and prototyping through implementation, production rollout, and ongoing operation. Your impact will come from the systems you build, the problems you solve, and measurable improvements in reliability, performance, and cost efficiency.
The technology stack runs on AWS, with workloads orchestrated on Kubernetes and data anchored in Aurora RDS (MySQL and Postgres). You will write code and infrastructure-as-code, debug production systems, and deliver changes that hold up under real traffic and failure conditions.
Key responsibilities include:
- Own the technical architecture and evolution of core infrastructure: identify system limits, prioritize technical improvements, and implement changes that support increasing traffic, data volume, and workload complexity
- Connect business understanding to improvements across the system: learn how key workflows behave in production, trace their impact across applications, data, and infrastructure, and partner across teams to improve performance, reliability, and cost efficiency
- Engineer for scale and performance: build capacity models, run load and stress tests, diagnose bottlenecks across compute, networking, Kubernetes, and databases
- Design and build AWS infrastructure: implement resilient account, IAM, network, and service architectures; address service quotas, fault isolation, and multi-AZ or multi-region requirements
- Build and operate the Kubernetes platform: improve cluster architecture, lifecycle automation, workload isolation, resource allocation, autoscaling, safe upgrades, and deployment reliability
- Scale and optimize Aurora RDS for MySQL and Postgres: tune queries and indexes, address connection and replication bottlenecks, plan capacity, improve failover behavior, and implement safe schema changes
- Improve reliability through engineering: define and instrument service-level objectives and error budgets; implement failure isolation, backpressure, load shedding, and safe retry behavior
- Participate in on-call rotation and drive technical recovery during serious incidents: use evidence-based triage, execute mitigations, communicate findings, and implement corrective actions from postmortems
- Implement and test disaster recovery: design backup, restore, and failover mechanisms against agreed recovery time and recovery point objectives
- Build infrastructure-as-code and operational automation: make provisioning, configuration, deployments, upgrades, and recovery reproducible, reviewed, testable, and recoverable
- Build observability that makes production diagnosable: improve metrics, logs, traces, dashboards, and actionable alerts
- Deliver safe infrastructure migrations: design phased rollouts, compatibility checks, validation, and rollback paths for changes to shared production systems
- Improve cloud cost efficiency through technical changes: right-size resources, improve utilization, tune autoscaling and storage, and quantify savings
- Build and evaluate AI-assisted operational tooling: apply AI to investigation, runbooks, anomaly analysis, and repetitive operations, with bounded permissions and reviewable actions
What We're Looking For
Required qualifications:
- Bachelor's degree in Computer Science or a similar technical field
- 12+ years of experience across infrastructure, platform, site reliability, software development, or related engineering disciplines, with substantial hands-on depth designing and operating production infrastructure at scale
- Deep expertise with AWS: production experience across compute, IAM, multi-account architectures, and networking, including VPC design and private connectivity
- Deep expertise with Kubernetes in production: cluster lifecycle, scheduling, resource management, autoscaling, networking, and troubleshooting business-critical workloads. EKS experience is strongly preferred
- Deep expertise with relational databases at scale, specifically RDS/Aurora (MySQL and/or Postgres): query performance, indexing, connection management, replication, high availability, failover, and verified backup and recovery
- A track record of personally delivering infrastructure scaling improvements: identifying constraints, measuring baseline behavior, implementing changes, and demonstrating gains in capacity, latency, reliability, or cost efficiency
- Strong coding and automation skills, using Golang, Python, or similar languages to build production tooling and eliminate operational toil, alongside infrastructure-as-code experience with Terraform or equivalent
- Strong systems fundamentals: Linux, networking, DNS, TLS, storage, concurrency, and distributed system failure modes, with the ability to debug problems across infrastructure and application boundaries
- Experience operating a 24/7, high-availability platform where downtime has direct customer or revenue impact, including hands-on incident response and postmortem remediation
- Willingness to participate in an on-call rotation and demonstrated ability to recover production systems under pressure using evidence-based triage, safe mitigation, and clear technical communication
- Experience implementing and testing disaster recovery against defined recovery objectives, including restoring data and validating service recovery
- Practical experience with observability, load testing, capacity planning, and safe CI/CD practices for shared production infrastructure
- Active use of AI tooling in engineering or operations, with practical judgment about its limitations and how to verify generated code, recommendations, and operational actions
- Ability to carry ambiguous technical problems from investigation through production delivery and collaborate across engineering disciplines to resolve system-wide constraints
Preferred qualifications:
- Experience in fintech, payments, or banking, operating infrastructure with demanding reliability, security, and audit requirements
- Experience with multi-region architectures, chaos engineering, and failure testing, including the consistency and recovery tradeoffs of distributed data systems
- Proficiency with Prometheus, Grafana, Loki, Tempo, or comparable observability systems
- Experience building internal platform capabilities and self-service tooling, including deployment automation, progressive delivery, and reusable infrastructure components
- Experience building AI-assisted incident investigation or operational automation with restricted access, auditable execution, and clear human review points
How We Work With You
Our talent team at Get A Job.ai carefully screens all candidates before submitting qualified professionals to our client. When you apply through our platform at getajob.ai, one of our recruiters will conduct an initial screening to understand your background and ensure alignment with the role requirements.
Once we confirm you're a strong match, we will submit your profile to the client for their consideration. Throughout the interview process, we remain your advocate and point of contact. Please do not contact the client directly, as all communication should flow through Get A Job.ai to ensure a smooth and professional process for everyone involved.
This is a remote position, and the role requires participation in an on-call rotation to support a high-availability production environment.
Pay
Compensation information for this role was not provided to Get A Job.ai. Our recruiting team will discuss compensation details during the screening process based on your experience level and location.
Equal Employment Opportunity: Get A Job.ai is committed to fair and equitable recruiting practices. We welcome applications from all qualified candidates regardless of race, color, religion, sex, national origin, age, disability, or any other protected characteristic.
Apply with Get A Job.ai
A recruiter will review your profile and submit you to the client. Do not contact the client directly.
Apply through Get A Job.ai. A recruiter will review your profile and submit you. Do not contact the client directly.
Apply through Get A Job.ai. A recruiter will review your profile and submit you.
Terms used in this posting
- on-call
- You may be required to be reachable and available to work outside normal scheduled hours, typically for a set rotation.
Explore Get A Job.ai online
About this role & career path
Traits that fit this role
- Adaptability
- Innovation
- Intellectual Curiosity
- Cautiousness
- Attention to Detail
Source: O*NET Work Styles (Distinctiveness Rank).
Typical preparation needed: Job Zone 3: Medium Preparation Needed. Most occupations in this zone require training in vocational schools, related on-the-job experience, or an associate's degree. — via O*NET
Industry news
- I was a software engineer who couldn't get excited about AI. Now I'm studying to be a nurse. - Business Insider
- ‘AI code apocalypse’ hasn’t hit engineers as hard as the industry may think - IT Brew
- AI tools have sparked a coding revolution. Software engineers are figuring out what comes next. - Business Insider
Source: O*NET (public-domain bulk data)
Salary & compensation
Workers in Computer & Mathematical occupations earn a national median of $98,769 — via US Census ACS / Data USA
Job details above are provided by the employer/source. The sections on this page are compiled from public data sources with AI assistance.
Accommodations: if you need a workplace accommodation to apply for or perform this job, see ADA.gov or EEOC.gov for guidance on your rights and how to request one.
Listing facts
- Role Principal Infrastructure Engineer
- Employer Get A Job.ai
- Location Remote · Remote-friendly
- Type Full Time
- Pay (from listing) Pay not listed
- Posted September 19, 2026
- Apply by October 19, 2026
- Overview Full job description on this page (1,033 words)
Facts above come from this job record on Get A Job.AI — not copied from third-party review sites.
Typical work in Infrastructure Engineer
Independent occupational context from O*NET (U.S. public-domain labor data). This is about the occupation, not a rewrite of this employer's posting.
- Consult with users, administrators, and engineers to identify business and technical requirements for proposed system modifications or technology purchases.
- Implement system renovation projects in collaboration with technical staff, engineering consultants, installers, and vendors.
- Keep abreast of changes in industry practices and emerging telecommunications technology by reviewing current literature, talking with colleagues, participating in educational programs, attending meetings or workshops, or participating in professional organizations or conferences.
- Review and evaluate requests from engineers, managers, and technicians for system modifications.
- Assess existing facilities' needs for new or modified telecommunications systems.
- Develop, maintain, or implement telecommunications disaster recovery plans to ensure business continuity.
Limited public data for this employer
We only show facts we can ground in public sources (Wikidata, O*NET, news/discussion links, or this listing). We do not invent Glassdoor-style ratings, salaries, or testimonials when data is thin. Use the listing facts, occupation context, and related openings below while we continue researching.
Explore related openings
Occupation family: Infrastructure Engineer
Keep exploring on Get A Job.ai
Not quite the right fit? Your next opportunity is a click away.
- Browse all jobs
- More jobs by category
- Remote jobs you can do from anywhere
- Research typical pay for this role
- Set a job alert so new matches reach you first
- Upload your resume to apply faster
Hiring instead? Post a job and reach candidates searching right now.
