The Repairs Lead will design and oversee the end‑to‑end hardware repair program for Anthropic’s expanding fleet of data centers. Responsibilities include setting global repair strategy, managing turnaround time and backlog, developing procedures and training, coordinating vendor and reverse‑logistics programs, sizing spares pools, analyzing failure patterns, leading cadence meetings, and communicating constraints to leadership. The role requires deep technical knowledge of server, network, optics, and GPU/accelerator hardware, experience with large‑scale break‑fix operations, vendor management, and data‑driven decision making in a remote‑friendly environment.
Data Center Global Repairs Program Support na Anthropic
Remoto - United States
Ver mais vagas na AnthropicSalary
USD 320,000 - 405,000
Requirements
Skills
- 8+ years of experience in data center operations as a manager, technical lead or related role, including accountability for production availability
- Proven track record running break‑fix programs at a large scale across multiple sites
- Managed vendors, OEMs, or contract workforces to measurable outcomes such as SLAs, operational reviews, and corrective action
- Hands‑on technical depth in server, network, and rack‑level hardware to independently verify repair quality and audit vendor claims
- Built or substantially improved operational processes
- Comfortable working with ticket, telemetry, and inventory data to drive decisions
- Bachelor’s degree in a relevant domain or equivalent practical experience
- Experience with GPU/accelerator or high‑density liquid‑cooled infrastructure, including tray, cold plate, and manifold‑level repair
- Experience managing RMA and warranty programs with hyperscale OEMs/ODMs, including failure analysis and supplier quality engagement
- Experience with spares planning, reverse logistics, or depot repair at data center scale
- Experience delivering repair outcomes inside partner‑operated or colocation sites where on‑floor operations are staffed via third parties
- Familiarity with optics and high‑speed interconnect failures
Responsibilities
- Define the global repair strategy including repair SLAs, prioritization rules, escalation paths, and reporting methods, and drive standardization across all sites
- Own repair turnaround time and repair backlog across the fleet evidenced by Anthropic‑owned ticket and telemetry dashboards you help develop
- Author and improve procedures for triage, break‑fix, return‑to‑service validation, and train site operations partners on how to execute them
- Manage RMA and reverse logistics programs with OEMs, ODMs, and depot repair vendors, including warranty claims, return cycle times, and failure analysis feedback
- Set spares pool sizing and stocking levels by site and part, in coordination with supply chain and asset management, so that parts availability never gates repair SLAs
- Analyze failure patterns across sites to identify root causes and drive corrective actions with hardware engineering, suppliers, and site operations owners
- Lead the operating cadence with vendor and site leads, including weekly repair reviews, scorecards, and business reviews, and drive corrective actions for SLA excursions
- Communicate repair constraints, risks, and fleet availability impact to engineering and leadership
Technologies
Server hardwareNetwork hardwareGPU/accelerator hardwareOptics and high‑speed interconnectRMA and reverse logistics processesOEMs and ODMsSpares inventory managementTicketing systemsTelemetry dashboards
Descubra se seu currículo está pronto para esta vaga
Veja como nossa IA pode otimizar seu currículo e aumentar suas chances de conseguir esta posição.