Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. The Scheduler team owns the scheduling layer for Anthropic's Kubernetes fleet, building the scheduling platform, job‑launch tooling, and fleet‑efficiency systems to ensure compute resources are used efficiently. We are looking for an engineering manager to lead this team, set technical direction, partner across capacity planning, research, inference, and product teams, and drive roadmap for scheduling capabilities and fleet utilization.
Engineering Manager, Scheduler and Fleet Efficiency en Anthropic
Híbrido - San Francisco, CA, United States
Más vacantes en AnthropicSalary
USD 405,000 - 485,000
Requirements
Skills
- Experience managing and growing a team of software engineers
- Hands‑on software engineering background as an individual contributor
- Experience building or operating large‑scale distributed or infrastructure systems in production
- Working knowledge of Kubernetes and cluster scheduling concepts such as resource requests and limits, affinity, priority and preemption, and custom schedulers or controllers
- Excellent written and verbal communication skills
- 5+ years of engineering management experience, including leading infrastructure, platform, or compute teams
- Experience owning a cluster scheduler, job orchestration system, or resource manager at scale
- Familiarity with scheduling ML workloads on accelerators and the tradeoffs between utilization, fairness, and latency
- Experience building developer tooling that other engineers rely on every day
- Background in observability or incident response for control‑plane systems, and a track record of improving production reliability
- Track record of building a culture of belonging and engineering excellence
- Low ego, high empathy, and habit of leading by example
Responsibilities
- Lead and grow a team of engineers building Anthropic's scheduling platform, job‑launch tooling, and fleet‑efficiency systems
- Set technical direction for scheduling, placement, queueing, and quota across Anthropic's compute fleet
- Partner with capacity planning, research, inference, and product teams to bring workloads onto the paved path and make efficient scheduling decisions
- Drive the roadmap for scheduler capabilities, fleet utilization, and the developer experience of launching and managing jobs
- Define and track metrics that measure fleet efficiency and scheduling quality such as utilization, queue wait, job‑start latency, etc., and hold the team accountable
- Create clarity for the team and stakeholders in an ambiguous, fast‑moving environment where demand for compute routinely exceeds supply
- Take an inclusive, equitable approach to hiring, coaching, and career development, and sustain a high‑performing, healthy team
- Represent the team across the engineering organization and contribute to engineering‑wide initiatives as a member of Anthropic's engineering management group
Technologies
KubernetesCluster schedulingCustom schedulers or controllersJob orchestration systemObservabilityIncident response for control‑plane systems
Descubre si tu currículum está listo para esta vacante
Mira cómo nuestra IA puede optimizar tu currículum y aumentar tus chances en este puesto.