Anthropic is at the forefront of AI research, dedicated to developing safe, ethical, and powerful artificial intelligence. The role involves building the infrastructure that measures what our models can actually do, ensuring evaluations are fast, reliable, and trustworthy at scale. Working at the intersection of inference, research, and infrastructure engineering, you will manage large‑scale distributed systems that orchestrate evaluations for frontier models, build and scale harnesses researchers use to run evals, make results reproducible and interpretable, and ensure eval signals reach decision points. Your work directly shapes what we build and what we don’t.
Evals Infrastructure Tech Lead / Manager en Anthropic
Híbrido - San Francisco, CA, United States
Más vacantes en AnthropicSalary
USD 500,000 - 850,000
Requirements
Skills
- Lead technical projects end-to-end on large‑scale distributed systems
- 1+ years managing engineers or tech‑lead-with-reports experience
- Strong proficiency in Python and Rust
- Built high‑throughput, fault‑tolerant systems on cloud or on‑prem accelerator fleets
- Care about measurement quality, not just pipeline uptime
- Communicate effectively with researchers and translate research needs into infrastructure
- Deeply interested in the transformative effects of advanced AI and committed to safe development
- Worked on LLM inference or training infrastructure
- Experience with eval or benchmarking systems, especially agentic evals requiring sandboxed execution
- Statistical literacy: variance, confidence intervals, sample‑size sufficiency for noisy metrics
- Experience with observability and regression detection over time‑series metrics
- Bachelor’s degree or an equivalent combination of education, training, and/or experience
Responsibilities
- Lead the team building the distributed systems that schedule, orchestrate, and execute evals for our frontier model training
- Own eval throughput and cost: compute allocation across suites, queueing against constrained accelerator pools, caching and reuse of eval work
- Build and scale the harnesses researchers use to define, run, and iterate on evals
- Make eval results trustworthy — determinism, reproducibility, and honest uncertainty quantification on reported metrics
- Ensure eval signal reaches the dashboards and reviews where launch decisions actually get made
- Contribute directly as an engineer while managing and growing the team, prioritizing its work, and coaching your reports
Technologies
PythonRustCloud platforms (e.g., AWS, GCP)On‑prem accelerator fleetsDistributed systemsSandboxed execution infrastructureObservability and regression detection tools
Descubre si tu currículum está listo para esta vacante
Mira cómo nuestra IA puede optimizar tu currículum y aumentar tus chances en este puesto.