Job Requirements
Education: Bachelor’s degree or above.
Work Experience: Minimum 5‑years of server and Linux O&M experience, including at least 3‑years of hands‑on experience in GPU compute‑cluster or AIDC production environments. Practical background in GPU‑server deployment, hardware troubleshooting, Linux administration and compute‑node resource management. Prior experience maintaining GPU server fleets of 100+ nodes or equivalent‑scale HPC clusters is preferred.
Core Skills:
- Familiar with GPU‑server architectures such as B300, GB200, HGX or DGX; capable of diagnosing hardware faults on GPU, NVLink/NVSwitch, NICs and memory modules.
- Proficient with BMC/IPMI/Redfish, BIOS, RAID and firmware administration.
- Expert in Rocky Linux / Ubuntu deployment, configuration, upgrade and troubleshooting.
- Knowledgeable about NVIDIA drivers, CUDA and Fabric Manager.
- Basic familiarity with storage systems including Ceph, Lustre and BeeGFS.
- Understand Kubernetes container orchestration and GPU resource scheduling.
- Experienced with monitoring stacks: Prometheus, Grafana, DCGM Exporter. Able to build automation workflows using Shell, Python or Ansible.
Certification Requirements: Must hold a valid RHCE certification. Candidates with CKA, CKS or NVIDIA AI Infrastructure certifications are preferred.
Other Special Requirements: Must be legally authorized to work in the United States. Able to work in both English and Chinese. Permitted to access computer rooms for cabling and on‑site troubleshooting.Key Job Responsibilities
- Complete GPU‑server receiving inspection, rack mounting, cabling, initialization, cluster onboarding and asset management.
- Diagnose hardware faults for GPU, NVLink/NVSwitch, memory and disks; coordinate component replacement and vendor RMA workflows.
- Manage full lifecycle for BMC, BIOS, firmware, Linux OS, NVIDIA drivers and CUDA.
- Conduct Linux configuration, patching, security hardening, log analysis and fault resolution.
- Perform day‑to‑day O&M, capacity planning and performance monitoring for Ceph, Lustre, BeeGFS storage systems.
- Collaborate with Shenzhen‑based teams on Kubernetes, Slurm, container operations, GPU scheduling and cluster incident resolution.
- Manage full lifecycle of servers and GPU nodes, and maintain compute‑resource inventory records.
- Coordinate with colocation facilities, server vendors and Shenzhen technical teams for emergency incident handling.
Pay: $65,000.00 - $90,000.00 per year
Language:
Work Location: In person