Loading...

Distributed Systems Engineer

  • Company: Get A Job.ai
  • Location: United States
  • Salary: Pay not listed

Website Get A Job.ai

Represented by Get A Job.ai

About This Opportunity

We are representing a confidential AI infrastructure organization operating in the United States. Our client is building distributed systems and operational infrastructure that support AI and exascale computing environments at significant scale. We're seeking a Distributed Systems Engineer to join their remote engineering team.

This role focuses on systems programming, distributed architecture, fault tolerance, and HPC-grade reliability. You'll be working on foundational infrastructure that powers large-scale AI compute environments.

Responsibilities

  • Design and build distributed systems that tolerate latency, bandwidth constraints, and intermittent connectivity
  • Implement fault-tolerant communication strategies, retry logic, backpressure, caching, and eventual consistency patterns
  • Write maintainable, resilient, and tested code following established development standards and methodologies
  • Debug and improve system behavior, including networking and distributed coordination issues
  • Contribute to systems written primarily in Rust
  • Work with system-level concerns including scheduling, memory management, I/O optimization, storage hierarchy management, and system reliability
  • Optimize performance and memory usage in resource-constrained environments
  • Debug concurrency issues and distributed coordination challenges
  • Design and maintain APIs and communication layers between distributed components
  • Identify and reduce tight coupling across services and systems
  • Diagnose and resolve cross-system failures in production environments
  • Collaborate with engineers and computer scientists on operating systems internals, compiler internals, fault tolerance, file system architecture, and trusted systems
  • Contribute to resolving architectural and systemic issues

What We're Looking For

Required:

  • Experience building distributed systems in environments with low bandwidth, high latency, or unreliable communication links
  • Production experience developing systems in Rust or Go
  • Understanding of distributed systems failure modes and mitigation strategies
  • Knowledge of consistency models, coordination strategies, and state replication
  • Experience designing APIs and communication layers between distributed components
  • Experience working within established architectures and delivering production-quality components
  • Understanding of systems-level concepts including durability, reliability, and operational behavior
  • Ability to work independently while collaborating with technical leadership

Preferred:

  • Experience with HPC environments, exascale computing, or AI/ML infrastructure
  • Exposure to operating systems internals, compiler design, or language runtimes
  • Experience with edge computing or constrained network environments
  • Familiarity with message queues, event-driven systems, or streaming architectures
  • Exposure to consensus algorithms or distributed coordination primitives
  • Experience with concurrency, memory management, or performance optimization in production systems
  • Experience contributing to developer tooling, internal platforms, or infrastructure-layer components

How We Work With You

Candidates apply directly through Get A Job.ai. Our recruiting team will conduct an initial screening to understand your background and ensure alignment with the role requirements. Qualified candidates will be submitted to our client for consideration. Please do not attempt to contact the client directly, as all communications are managed through our talent team to ensure a smooth process for everyone involved.

This is a remote, full-time exempt position based in the United States. You'll have the opportunity to work on distributed systems supporting AI and exascale workloads, collaborating with engineers experienced in low-level systems, fault tolerance, and large-scale infrastructure.

Pay

Compensation details will be discussed with qualified candidates during the screening process.

Equal Opportunity: Get A Job.ai is committed to inclusive recruiting practices. We work with clients who value diversity and provide equal employment opportunities to all qualified applicants.

Apply with Get A Job.ai

A recruiter will review your profile and submit you to the client. Do not contact the client directly.

More options

Apply with Get A Job.ai

Apply through Get A Job.ai. A recruiter will review your profile and submit you. Do not contact the client directly.

Apply through Get A Job.ai. A recruiter will review your profile and submit you.

Local insights for this role are preparing — this section updates automatically in a few seconds (or refresh).

Listing facts

  • Role Distributed Systems Engineer
  • Employer Get A Job.ai
  • Location United States
  • Type Full Time
  • Pay (from listing) Pay not listed
  • Posted September 2, 2026
  • Apply by October 2, 2026
  • Country United States
  • Overview Full job description on this page (511 words)

Facts above come from this job record on Get A Job.AI — not copied from third-party review sites.

Limited public data for this employer

We only show facts we can ground in public sources (Wikidata, O*NET, news/discussion links, or this listing). We do not invent Glassdoor-style ratings, salaries, or testimonials when data is thin. Use the listing facts, occupation context, and related openings below while we continue researching.

Explore related openings

Occupation family: Distributed Systems Engineer

Keep exploring on Get A Job.ai

Not quite the right fit? Your next opportunity is a click away.

Hiring instead? Post a job and reach candidates searching right now.

Distributed Systems Engineer Get A Job.ai · United States