Lead, Hardware Deployment Engineer na xAI

Remoto - Memphis, Tennessee, United States

Candidatar-se
Ver mais vagas na xAI

SpaceXAI seeks a Lead Hardware Deployment Engineer to oversee end‑to‑end GPU compute hardware bring‑up across its AI training clusters. The role involves managing a deployment team, driving integration and repair processes, and ensuring high availability of hardware across multiple data halls. The position is based in Memphis, Tennessee, with remote flexibility.

Requirements

Skills

  • 5+ years of hands‑on experience deploying, integrating, or repairing compute/server hardware at data center scale.
  • Direct experience with L11 (rack‑level) integration and bring‑up of GPU or accelerator‑based systems.
  • Demonstrated experience leading technician or engineering teams in a fast‑paced deployment, manufacturing, or data center environment.
  • Deep troubleshooting skills across servers, GPUs, NVLink/fabric interconnects, high‑speed networking, and liquid cooling systems.
  • Willingness to work on‑site in Memphis, TN, including extended hours and weekends during critical bring‑up phases.

Responsibilities

  • Lead, hire, and develop a dedicated hardware deployment team (deployment engineers, deployment technicians, and repair technicians) with full ownership of team structure and staffing.
  • Own L11 rack integration and compute hardware bring‑up across multiple data halls concurrently, from delivery dock to healthy production handoff.
  • Drive aggressive bring‑up timelines: achieve 95%+ node availability within days of rack delivery and 100% closure within one week per data hall.
  • Own post‑L11 hardware health: run systematic health pushes to sustain greater than 98% node availability prior to turnover to operations.
  • Internalize non‑RMA hardware repairs to maximize hardware recovery, minimize repair backlogs, and reduce dependence on OEM turnaround times.
  • Develop and enforce vendor SLAs for OEM and supplier responsibilities; prevent accumulation of unrepaired hardware ("bone piles") and repair backlogs before turnover to operations.
  • Perform root cause analysis of hardware failures discovered during L11 and drive corrective actions with vendors and internal engineering teams.
  • Partner with site operations on hardware debugging and repair, and train site operations teams to support future data center deployments.
  • Build, document, and continuously improve deployment processes, tooling, and training so bring‑up capability scales across sites and future hardware generations.

Technologies

NVIDIA GPUNVLinkLiquid cooling system

Compartilhar vaga

Descubra se seu currículo está pronto para esta vaga

Veja como nossa IA pode otimizar seu currículo e aumentar suas chances de conseguir esta posição.