This role at Anthropic focuses on Capacity Engineering, responsible for building production systems that monitor and optimize infrastructure utilization across multiple cloud providers. The engineer will design data pipelines, observability tooling, and performance instrumentation to enable efficient allocation and cost attribution at scale.
Staff+ Software Engineer, Capacity Engineering at Anthropic
On-site - San Francisco, CA; New York City, NY; Seattle, WA
More jobs at AnthropicRequirements
Skills
- Strong track record building and operating production systems
- Hands‑on engineering role with devops flavor
- Python at production quality
- SQL at production quality
- Deep experience with at least one major cloud provider (Amazon Web Services, Google Cloud, or Microsoft Azure) and its operations
- Experience with observability tooling stack, including Prometheus, PromQL, and Grafana, including writing recording rules and building monitoring that engineering teams rely on
- Ability to gather your own requirements and work across organizational boundaries in an ambiguous environment
Responsibilities
- Build the planning and allocation stack—tools leadership uses to allocate capacity, teams use to plan against their allocations, and the scheduler enforces; cross‑region and cross‑provider placement, guardrails, queueing, occupancy KPIs
- Drive the efficiency programs: stranding and rightsizing, unused capacity recovery, and job‑level utilization across training, inference, and eval; establish per‑config baselines and work with system‑owning teams to close the gaps
- Own attribution and forecasting—reconcile billing across ten‑plus providers against telemetry and internal systems, attribute spend to the workloads that generate it, and turn demand signals and research roadmaps into a defensible compute plan and supply pipeline
- Build the data platform underneath all of it: pipelines ingesting occupancy, utilization, and cost from a rapidly diversifying fleet into BigQuery, with real ownership of completeness, latency SLOs, and gap detection
- Operate Kubernetes‑native systems at scale—collection agents, workload labeling, and the taint/reservation/scheduling behavior that determines what capacity is actually usable
- Treat the output as a product, not a pipeline; gather your own requirements, define schema contracts, and design for consumers ranging from research engineers to a CFO—including on‑call and SLOs
Technologies
PythonSQLBigQueryPrometheusPromQLGrafanaKubernetesAmazon Web ServicesGoogle CloudMicrosoft Azure
See if your resume is ready for this job
See how our AI can optimize your resume and improve your chances for this role.