Engineering Manager, Scheduler and Fleet Efficiency na Anthropic

Híbrido - San Francisco, CA, United States

Candidatar-se
Ver mais vagas na Anthropic

Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. The Scheduler team owns the scheduling layer for Anthropic's Kubernetes fleet, building the scheduling platform, job‑launch tooling, and fleet‑efficiency systems to ensure compute resources are used efficiently. We are looking for an engineering manager to lead this team, set technical direction, partner across capacity planning, research, inference, and product teams, and drive roadmap for scheduling capabilities and fleet utilization.

Salary

USD 405,000 - 485,000

Requirements

Skills

  • Experience managing and growing a team of software engineers
  • Hands‑on software engineering background as an individual contributor
  • Experience building or operating large‑scale distributed or infrastructure systems in production
  • Working knowledge of Kubernetes and cluster scheduling concepts such as resource requests and limits, affinity, priority and preemption, and custom schedulers or controllers
  • Excellent written and verbal communication skills
  • 5+ years of engineering management experience, including leading infrastructure, platform, or compute teams
  • Experience owning a cluster scheduler, job orchestration system, or resource manager at scale
  • Familiarity with scheduling ML workloads on accelerators and the tradeoffs between utilization, fairness, and latency
  • Experience building developer tooling that other engineers rely on every day
  • Background in observability or incident response for control‑plane systems, and a track record of improving production reliability
  • Track record of building a culture of belonging and engineering excellence
  • Low ego, high empathy, and habit of leading by example

Responsibilities

  • Lead and grow a team of engineers building Anthropic's scheduling platform, job‑launch tooling, and fleet‑efficiency systems
  • Set technical direction for scheduling, placement, queueing, and quota across Anthropic's compute fleet
  • Partner with capacity planning, research, inference, and product teams to bring workloads onto the paved path and make efficient scheduling decisions
  • Drive the roadmap for scheduler capabilities, fleet utilization, and the developer experience of launching and managing jobs
  • Define and track metrics that measure fleet efficiency and scheduling quality such as utilization, queue wait, job‑start latency, etc., and hold the team accountable
  • Create clarity for the team and stakeholders in an ambiguous, fast‑moving environment where demand for compute routinely exceeds supply
  • Take an inclusive, equitable approach to hiring, coaching, and career development, and sustain a high‑performing, healthy team
  • Represent the team across the engineering organization and contribute to engineering‑wide initiatives as a member of Anthropic's engineering management group

Technologies

KubernetesCluster schedulingCustom schedulers or controllersJob orchestration systemObservabilityIncident response for control‑plane systems

Compartilhar vaga

Descubra se seu currículo está pronto para esta vaga

Veja como nossa IA pode otimizar seu currículo e aumentar suas chances de conseguir esta posição.