JOB SUMMARY
We are seeking an experienced RMA Failure Analysis for GPU Servers and enterprise server platforms. The engineer will be responsible for diagnosing, troubleshooting, and performing root cause analysis on customer-returned GPU servers, server motherboards, GPU baseboards, and associated hardware subsystems. This ideal candidate will possess strong server architecture knowledge, component-level debugging expertise, and the ability to safely handle and analyze high-value hardware throughout the failure analysis process.
Key Responsibilities
- Perform failure analysis on customer-returned GPU servers, server motherboards, GPU boards, GPU baseboards, and related hardware assemblies.
- Conduct system-level, board-level, and component-level troubleshooting to identify root causes of hardware failures.
- Execute functional testing, diagnostics, and debug activities using standard lab equipment and server validation tools.
- Read and interpret schematics, block diagrams, board layouts, and manufacturing documentation.
- Analyze failures involving server subsystems including CPUs, GPUs, DIMMs, NICs, SSDs, power supplies, PCIe devices, and cooling/thermal subsystems.
- Troubleshoot hardware issues related to BIOS, BMC, CPLD, FPGA, PCIe, memory, storage, networking, and power delivery circuits.
- Perform component-level debugging including capacitors, resistors, fuses, diodes, MOSFETs, voltage regulators, ICs, and other electronic components.
- Conduct component swapping, isolation testing, and fault reproduction to validate failure mechanisms and root causes.
- Perform detailed visual and mechanical inspections to identify damaged, missing, misaligned, overheated, or improperly assembled components.
- Utilize JIRA and Zendesk to track RMA cases, document failure analysis results, manage issue resolution activities, and maintain clear communication across engineering, quality, and customer support teams.
- Document failure analysis findings, corrective actions, and recommendations to support continuous product quality improvements.
- Collaborate with design, validation, manufacturing, and quality teams to drive issue resolution and corrective actions.
- Follow proper ESD and hardware handling procedures while working with customer-returned products, engineering samples, and production hardware.
Required Qualifications
- 4+ years of experience in server hardware design, validation, testing, debugging, failure analysis, or system engineering.
- Strong understanding of GPU server architecture and enterprise server platforms.
- Experience performing system-level, board-level, and component-level troubleshooting.
- Ability to read and interpret electrical schematics, block diagrams, and PCB layouts.
- Hands-on experience with server technologies including BIOS, BMC, CPLD, FPGA, PCIe, memory subsystems, storage interfaces, and networking interfaces.
- Experience using laboratory equipment such as oscilloscopes, digital multimeters (DMM), power analyzers, logic analyzers, and protocol analyzers.
- Working knowledge of Linux operating systems and command-line troubleshooting.
- Strong understanding of root cause analysis methodologies and failure isolation techniques.
- Ability to safely handle sensitive server and GPU hardware while adhering to ESD and hardware handling best practices.
Preferred Qualifications
- Experience supporting AI, HPC, or GPU-accelerated server platforms.
- Experience with customer-returned hardware (RMA) failure analysis processes.
- Knowledge of power delivery architecture, thermal analysis, and signal integrity concepts.
- Familiarity with manufacturing defects, field failures, and reliability-related investigations.
Critical Requirements
- Must be capable of independently troubleshooting GPU servers and server hardware down to the component level.
- Must understand overall server architecture and subsystem interactions before initiating debug activities.
- Must demonstrate strong analytical and problem-solving skills in hardware failure analysis.
- Must be comfortable working with customer-returned hardware and managing multiple RMA investigations simultaneously.
- Must maintain proper hardware handling practices to prevent damage to customer-returned units and engineering samples.
Pay: $25.00 - $30.00 per hour
Work Location: In person