Key Responsibilities
- Execute and coordinate rack-level deployment of compute, network, and storage infrastructure for AI and HPC environments.
- Validate rack elevations, port maps, cable maps, power distribution, labeling, and physical connectivity before cluster handoff.
- Perform detailed cable validation across Ethernet, InfiniBand, management, and storage interconnects.
- Detect and resolve cabling defects including polarity issues, incorrect port destinations, unsupported optics pairings, breakout errors, damaged media, and inconsistent labeling.
- Support hardware bring-up by validating BIOS baselines, BMC reachability, inventory accuracy, and initial connectivity tests.
- Partner with network and Linux teams during cluster turn-up to quickly isolate physical-layer versus logical-layer failures.
- Create deployment checklists, as-built documentation, and signoff criteria for production acceptance.
- Drive structured remediation during expansion, re-cabling, or failed build events.
Required Qualifications
- 7+ years in data center deployment, HPC infrastructure installation, or large-scale hardware integration.
- Strong practical knowledge of structured cabling, optics, transceivers, breakout schemes, and rack-level physical validation.
- Experience reading and validating rack diagrams, patch plans, cable matrices, and topology documentation.
- Strong troubleshooting skill for Layer 1 issues that manifest as network instability, missing hosts, degraded performance, or failed cluster readiness checks.
- Experience with asset tracking, labeling discipline, and deployment quality control.
- Ability to work across physical infrastructure, server hardware, and network operations teams without losing detail.
Pay: $90.00 - $95.00 per hour
Experience:
- HPC: 6 years (Required)
- data center: 8 years (Required)
Willingness to travel:
Work Location: Hybrid remote in Santa Clara, CA 95051