The Capacity Deployment Lead will manage and scale the deployment of compute capacity across partner‑operated data centers, overseeing construction, installation, QA, and handover processes while driving continuous improvement across the program and ensuring on‑site delivery standards.
Capacity Deployment Lead - Data Center Operations na Anthropic
Remoto - United States
Ver mais vagas na AnthropicSalary
USD 320,000 - 405,000
Requirements
Skills
- Have 8+ years of experience in data center operations or hardware deployment as a manager, technical lead or related role, including accountability for delivery milestones.
- Can demonstrate a proven track record leading compute or network hardware deployments at a large scale across multiple sites or data-hall waves, from early floor access through ready for service.
- Have managed vendors, OEMs, or contract workforces to measurable outcomes: milestones, quality gates, operational reviews, and corrective action.
- Possess hands‑on technical depth in server, network, and rack‑level hardware, enough to independently verify deployment quality and audit vendor claims.
- Have built or substantially improved operational processes.
- Are comfortable working with schedule, ticket, and telemetry data to drive decisions.
- Can travel heavily and work on site for extended stretches, including during delivery surges and turn‑up windows. 50% Travel expected.
- Possess a bachelor's degree in relevant domain or equivalent practical experience.
Responsibilities
- Own deployment outcomes for your assigned sites and data-hall waves: CUs and hall ready-for-service dates met, with exceptions closed out, or risk called early enough for leadership to act.
- Engage from construction kickoff and confirm deployment readiness ahead of each wave - floor access, power, network, receiving paths, and partner staffing - working the gaps with the site operations partner and facilities counterparts before racks arrive.
- Be on site at key points before and during delivery, deployment, QA, and close-out; verify the work with your own eyes rather than from a dashboard, and support the on-site teams with your direct expertise.
- Work with partner's site leadership to ensure success through rack delivery and receiving, integration QA, power-on and network turn‑up, and burn-in and validation, setting daily priorities and quality gates and confirming results against Anthropic-owned ticket and telemetry data.
- Validate the handover gate for each CU and data hall: acceptance criteria met, burn‑in complete, documentation delivered, spares positioned, ticketing live, and a structured handover into production operations and the maintenance program.
- Turn best practices, optimization, and lessons learned from each deployment into program-wide improvements: procedures, checklists, acceptance standards, and metrics, so the next wave runs faster and cleaner than the last.
- Extend the deployment program to new sites and platforms, adapting standards and acceptance criteria as the hardware changes.
- Communicate deployment status, dependencies, constraints, and risk to engineering, capacity planning, the team that owns the overall site schedule, and leadership.
Technologies
GPU/acceleratorhigh-density liquid-cooled infrastructureserver hardwarenetwork hardwarerack hardwareoptical cablinghigh-speed interconnectHPC
Descubra se seu currículo está pronto para esta vaga
Veja como nossa IA pode otimizar seu currículo e aumentar suas chances de conseguir esta posição.