Data Center Operations Lead - Partner Site Operations en Anthropic

Remoto - San Francisco, CA, USA

Postularse
Más vacantes en Anthropic

Anthropic’s Data Center Operations (DCO) team ensures compute fleet availability through hardware and IT operations. This role leads partner‑operated sites, providing tactical direction, defining standards, and overseeing vendor performance to achieve site availability, deployment velocity, and incident response goals.

Salary

USD 320,000 - 405,000

Requirements

Skills

  • 8+ years of experience in data center operations (hardware, IT infrastructure, or critical facilities) as a manager, technical lead or related role
  • Managed vendors, MSPs, or contract workforces to measurable outcomes (SOWs, SLAs, operational reviews, corrective action)
  • Hands‑on technical depth in server, network, and rack‑level infrastructure to verify vendor claims
  • Built or substantially improved operational processes
  • Experience in incident command or lead‑responder role with clear communication under ambiguity
  • Availability for non‑standard hours, on‑call rotation and deployment surges
  • Bachelor's degree in a relevant domain or equivalent practical experience

Responsibilities

  • Own site availability, deployment milestones, and repair turnaround
  • Set daily and weekly priorities and lead the operating cadence, including stand‑ups and business reviews
  • Define and improve procedures for deployment, break‑fix, change management, security, and EHS compliance
  • Track vendor performance against SLAs and staffing commitments, driving corrective actions
  • Participate in incident escalation on‑call rotation and serve as Incident Commander for site‑specific incidents
  • Translate engineering requirements into vendor direction and communicate site constraints and risks to leadership
  • Lead weekly operations reviews and scorecards with vendor site leads
  • Direct deployment surges to meet first‑compute‑online milestones
  • Analyze failure patterns to identify root causes and drive fixes
  • Create break‑fix ownership matrices and train vendor teams
  • Serve as Incident Commander for facility events and produce post‑mortems
  • Establish operational readiness for new data halls, including spares and security
  • Identify process gaps and codify improvements as program standards

Technologies

server infrastructurenetwork infrastructurerack‑level infrastructureGPU/accelerator infrastructurehigh‑density liquid‑cooled infrastructureEHS compliance tools

Compartir vacante

Descubre si tu currículum está listo para esta vacante

Mira cómo nuestra IA puede optimizar tu currículum y aumentar tus chances en este puesto.