Anthropic’s Data Center Operations (DCO) team ensures compute fleet availability through hardware and IT operations. This role leads partner‑operated sites, providing tactical direction, defining standards, and overseeing vendor performance to achieve site availability, deployment velocity, and incident response goals.
Data Center Operations Lead - Partner Site Operations na Anthropic
Remoto - San Francisco, CA, USA
Ver mais vagas na AnthropicSalary
USD 320,000 - 405,000
Requirements
Skills
- 8+ years of experience in data center operations (hardware, IT infrastructure, or critical facilities) as a manager, technical lead or related role
- Managed vendors, MSPs, or contract workforces to measurable outcomes (SOWs, SLAs, operational reviews, corrective action)
- Hands‑on technical depth in server, network, and rack‑level infrastructure to verify vendor claims
- Built or substantially improved operational processes
- Experience in incident command or lead‑responder role with clear communication under ambiguity
- Availability for non‑standard hours, on‑call rotation and deployment surges
- Bachelor's degree in a relevant domain or equivalent practical experience
Responsibilities
- Own site availability, deployment milestones, and repair turnaround
- Set daily and weekly priorities and lead the operating cadence, including stand‑ups and business reviews
- Define and improve procedures for deployment, break‑fix, change management, security, and EHS compliance
- Track vendor performance against SLAs and staffing commitments, driving corrective actions
- Participate in incident escalation on‑call rotation and serve as Incident Commander for site‑specific incidents
- Translate engineering requirements into vendor direction and communicate site constraints and risks to leadership
- Lead weekly operations reviews and scorecards with vendor site leads
- Direct deployment surges to meet first‑compute‑online milestones
- Analyze failure patterns to identify root causes and drive fixes
- Create break‑fix ownership matrices and train vendor teams
- Serve as Incident Commander for facility events and produce post‑mortems
- Establish operational readiness for new data halls, including spares and security
- Identify process gaps and codify improvements as program standards
Technologies
server infrastructurenetwork infrastructurerack‑level infrastructureGPU/accelerator infrastructurehigh‑density liquid‑cooled infrastructureEHS compliance tools
Descubra se seu currículo está pronto para esta vaga
Veja como nossa IA pode otimizar seu currículo e aumentar suas chances de conseguir esta posição.