Senior Engineering Manager, Capacity Engineering at Anthropic

On-site - San Francisco, CA, USA

Apply
More jobs at Anthropic

Anthropic’s capacity engineering team leads production systems ensuring efficient utilization of compute resources across multiple cloud providers and infrastructure. The Senior Engineering Manager will set technical direction, grow a senior and staff‑level engineering team, and oversee reliability, correctness, and operational excellence of capacity planning, data platform, and efficiency initiatives.

Requirements

Skills

  • Experience managing software or infrastructure engineering teams, including hiring senior engineers, managing performance, and developing people into larger scope.
  • A strong technical background in production systems — data engineering, infrastructure, distributed systems, or observability — with hands‑on experience you can still draw on when reviewing designs or debugging with the team.
  • Familiarity with at least one major cloud provider (AWS, GCP, or Azure), Kubernetes‑based infrastructure, and modern observability stacks (e.g., Prometheus, Grafana).
  • A track record of setting and executing an engineering roadmap in an ambiguous, high‑autonomy environment with many stakeholders and shifting priorities.
  • Excellent communication skills: you can explain a utilization metric to a research engineer and a spend forecast to a CFO, and you can advocate clearly for your team’s priorities with senior leadership.
  • Comfort owning operational responsibility for systems the company depends on, including on‑call and incident management.

Responsibilities

  • Be hands‑on, lead and grow the team: hire, onboard, coach, and retain senior and staff engineers; set expectations; give feedback; run performance and leveling conversations; build a culture of ownership, rigor, and collaboration.
  • Champion internal customers: engage directly with research engineering, inference, infrastructure, and finance teams; bring learnings back to the roadmap; lead the team in building tools people want to use.
  • Own the roadmap: translate compute strategy into an engineering roadmap across data platform, planning, and efficiency; make trade‑offs and communicate them.
  • Set the technical bar: review designs, weigh architecture, enforce production standards—well‑tested Python and SQL, SLOs, gap detection, and sustainable on‑call practices.
  • Run the team as a product organization: gather requirements, define schema contracts, design for diverse consumers.
  • Be the primary partner for cross‑functional stakeholders: work closely with infrastructure, inference, research engineering, and finance leadership to align on capacity decisions, efficiency targets, and spend.
  • Drive operational excellence: own reliability and incident response, establish SLOs and on‑call practices, reduce operational toil.
  • Scale the function: anticipate headcount, skill, and system growth; build case for expansion.

Technologies

PythonSQLPrometheusGrafanaKubernetesAWSGCPAzureBigQueryDCGM

See if your resume is ready for this job

See how our AI can optimize your resume and improve your chances for this role.