- Company: Get A Job.ai
- Location: United States
- Salary: Pay not listed
Website Get A Job.ai
Represented by Get A Job.ai
About This Opportunity
We are representing a confidential AI infrastructure organization operating in the United States. Our client is building distributed systems and operational infrastructure that support AI and exascale computing environments at significant scale. We're seeking a Distributed Systems Engineer to join their remote engineering team.
This role focuses on systems programming, distributed architecture, fault tolerance, and HPC-grade reliability. You'll be working on foundational infrastructure that powers large-scale AI compute environments.
Responsibilities
- Design and build distributed systems that tolerate latency, bandwidth constraints, and intermittent connectivity
- Implement fault-tolerant communication strategies, retry logic, backpressure, caching, and eventual consistency patterns
- Write maintainable, resilient, and tested code following established development standards and methodologies
- Debug and improve system behavior, including networking and distributed coordination issues
- Contribute to systems written primarily in Rust
- Work with system-level concerns including scheduling, memory management, I/O optimization, storage hierarchy management, and system reliability
- Optimize performance and memory usage in resource-constrained environments
- Debug concurrency issues and distributed coordination challenges
- Design and maintain APIs and communication layers between distributed components
- Identify and reduce tight coupling across services and systems
- Diagnose and resolve cross-system failures in production environments
- Collaborate with engineers and computer scientists on operating systems internals, compiler internals, fault tolerance, file system architecture, and trusted systems
- Contribute to resolving architectural and systemic issues
What We're Looking For
Required:
- Experience building distributed systems in environments with low bandwidth, high latency, or unreliable communication links
- Production experience developing systems in Rust or Go
- Understanding of distributed systems failure modes and mitigation strategies
- Knowledge of consistency models, coordination strategies, and state replication
- Experience designing APIs and communication layers between distributed components
- Experience working within established architectures and delivering production-quality components
- Understanding of systems-level concepts including durability, reliability, and operational behavior
- Ability to work independently while collaborating with technical leadership
Preferred:
- Experience with HPC environments, exascale computing, or AI/ML infrastructure
- Exposure to operating systems internals, compiler design, or language runtimes
- Experience with edge computing or constrained network environments
- Familiarity with message queues, event-driven systems, or streaming architectures
- Exposure to consensus algorithms or distributed coordination primitives
- Experience with concurrency, memory management, or performance optimization in production systems
- Experience contributing to developer tooling, internal platforms, or infrastructure-layer components
How We Work With You
Candidates apply directly through Get A Job.ai. Our recruiting team will conduct an initial screening to understand your background and ensure alignment with the role requirements. Qualified candidates will be submitted to our client for consideration. Please do not attempt to contact the client directly, as all communications are managed through our talent team to ensure a smooth process for everyone involved.
This is a remote, full-time exempt position based in the United States. You'll have the opportunity to work on distributed systems supporting AI and exascale workloads, collaborating with engineers experienced in low-level systems, fault tolerance, and large-scale infrastructure.
Pay
Compensation details will be discussed with qualified candidates during the screening process.
Equal Opportunity: Get A Job.ai is committed to inclusive recruiting practices. We work with clients who value diversity and provide equal employment opportunities to all qualified applicants.
Apply with Get A Job.ai
A recruiter will review your profile and submit you to the client. Do not contact the client directly.
Apply through Get A Job.ai. A recruiter will review your profile and submit you. Do not contact the client directly.
Apply through Get A Job.ai. A recruiter will review your profile and submit you.
Local insights for this role are preparing — this section updates automatically in a few seconds (or refresh).
Listing facts
- Role Distributed Systems Engineer
- Employer Get A Job.ai
- Location United States
- Type Full Time
- Pay (from listing) Pay not listed
- Posted September 2, 2026
- Apply by October 2, 2026
- Country United States
- Overview Full job description on this page (511 words)
Facts above come from this job record on Get A Job.AI — not copied from third-party review sites.
Limited public data for this employer
We only show facts we can ground in public sources (Wikidata, O*NET, news/discussion links, or this listing). We do not invent Glassdoor-style ratings, salaries, or testimonials when data is thin. Use the listing facts, occupation context, and related openings below while we continue researching.
Explore related openings
Occupation family: Distributed Systems Engineer
Keep exploring on Get A Job.ai
Not quite the right fit? Your next opportunity is a click away.
- Browse all jobs
- More jobs by category
- Remote jobs you can do from anywhere
- Research typical pay for this role
- Set a job alert so new matches reach you first
- Upload your resume to apply faster
Hiring instead? Post a job and reach candidates searching right now.
