KMC Solutions Inc
You will be working on validating and testing GPU clusters prior to production release, ensuring hardware integrity, system reliability, and optimal performance. This role involves provisioning clusters, executing performance benchmarks, maintaining automated validation frameworks, and troubleshooting Linux-based systems in high-performance compute environments. You will collaborate closely with engineering and operations teams to ensure seamless handovers and production readiness.
.• Health Insurance/HMO
• Enjoy unlimited MadMax Coffee
• Diverse learning & growth opportunities
• Accessible Cloud HR platform (Sprout)
• Above standard leaves
Cluster Validation & Testing
-
Validate GPU clusters of varying sizes to ensure hardware and system integrity prior to production release
-
Perform functional and reliability testing of GPUs, servers, and associated components
-
Verify network connectivity and performance, including InfiniBand where applicable
Orchestration & Benchmarking
-
Provision and configure GPU clusters using automated workflows
-
Execute and analyse performance and stability benchmarks orchestrated via Slurm
-
Validate results against expected performance and reliability thresholds
Test Framework & Automation
-
Maintain and extend the automated validation framework built using Python and Ansible
-
Integrate new test cases to support additional hardware platforms and GPU generations
-
Improve test reliability, coverage, and execution efficiency
Remediation & System Integrity
-
Diagnose and remediate unhealthy nodes through configuration changes or software fixes
-
Coordinate with on-site support and Smart Hands teams for hardware replacements when required
-
Ensure all issues are resolved and documented prior to handover to production operations
Documentation & Handover
-
Produce clear, accurate documentation of test results, hardware states, and remediation actions
-
Ensure smooth handovers to operations and engineering teams
-
Maintain up-to-date runbooks and validation procedures
Essential
• Strong hands-on experience administering and troubleshooting Linux systems (Prio)
• Confident use of CLI tools for diagnostics, including analysis of kernel logs, drivers, and system
services
• Excellent written and verbal English communication skills
• High standards for system reliability, consistency, and documentation
Preferred / Desirable
• Experience working with GPU-based or high-performance compute environments
• Familiarity with Slurm or other workload schedulers
• Understanding of datacenter hardware lifecycle and server validation processes
• Exposure to InfiniBand or high-speed networking technologies
• Experience working with distributed or remote infrastructure teams
• Proficiency in Python for automation, test execution, and parsing results (Preferred)
• Proven experience writing and maintaining Ansible playbooks (Preferred)
.
Originally posted on Himalayas
To apply for this job please visit himalayas.app.
Working in United States
The United States of America (USA), also known as the United States (U.S.) or America, is a country primarily located in North America. It is a federal republic consisting of 50 states and a federal capital district, Washington, D.C. The 48 contiguous states border Canada to the north and Mexico to the south, with the semi-exclave of Alaska in the northwest and the archipelago of Hawaii in the Pacific Ocean. The United States also asserts sovereignty over five major island territories and various uninhabited islands in Oceania and the Caribbean. It is a megadiverse country, with the world's th
More jobs at KMC Solutions Inc
Keep exploring on Get A Job.ai
Not quite the right fit? Your next opportunity is a click away.
- Browse all jobs
- More jobs by category
- Remote jobs you can do from anywhere
- Research typical pay for this role
- Set a job alert so new matches reach you first
- Upload your resume to apply faster
Hiring instead? Post a job and reach candidates searching right now.