The Hardware Systems Engineer is responsible for diagnosing, testing, repairing, and validating next-generation AI compute hardware supporting Oracle Cloud Infrastructure (OCI). As a core member of the Oracle Repair Center, this role ensures Compute hardware is rapidly returned to service while helping improve hardware reliability, diagnosability, and operational efficiency across the repair ecosystem. The position plays an important role in sustaining repair velocity and supporting the broader scale-up of Oracle’s AI infrastructure.
Working closely with Development, Manufacturing, Supply Chain, and Global Product Engineering, the Hardware Systems Engineer performs advanced hardware troubleshooting, root cause analysis, and physical repair of AI infrastructure, with a particular focus on NVIDIA and AMD compute hardware and future AI platforms. The role includes functional validation of Compute Trays (CTs) using standardized repair and test processes, execution of repair validation through dedicated CDU test infrastructure, and completion of engineering observation reports and repair documentation for manufacturing, quality, and engineering teams. In addition, the engineer serves as a technical resource during coolant leak events, supports continuous improvement efforts, and may participate in weekend coverage and on-call rotations for critical events.