Staff Research Engineer, Multi-Agent Scaling at Anthropic

Hybrid - San Francisco, CA, New York City, NY, Seattle, WA

Apply
More jobs at Anthropic

This role sits at the intersection of research and engineering within Anthropic’s AI Research & Engineering department, focusing on scaling multi‑agent systems. The team designs, runs, and interprets large‑scale experiments to understand how agent team performance, efficiency, and coordination change with scale. Responsibilities include building and scaling the infrastructure that runs these experiments, developing evaluation metrics, and collaborating with researchers across the organization. The position offers a hybrid work model with on‑site presence in one of Anthropic’s offices in San Francisco, New York City, or Seattle.

Salary

USD 500,000 - 850,000

Requirements

Skills

  • Significant software engineering, ML or research engineering experience
  • Owned something substantial end to end, such as a large system, an evaluation or benchmark, an agent product, or a research project
  • Genuinely enjoy both research and engineering work
  • Think quantitatively about complex systems, and think twice before trusting a number
  • Can work from a vague question rather than a spec
  • Results-oriented, with a bias towards flexibility and impact
  • Have clear written and verbal communication
  • Care about the societal impacts of your work
  • Experience building or operating large-scale distributed systems, such as schedulers, sandboxed code execution, or inference and RL infrastructure
  • Built evaluations, benchmarks or harnesses for LLMs or agents
  • Experience building complex agentic systems that use LLMs
  • Experience with scaling laws or other large-scale empirical research
  • A background in operations research, statistics, economics, physics, quantitative finance, or another field that models and optimizes complex systems

Responsibilities

  • Design, run and interpret large-scale experiments on agent teams, reasoning rigorously about what the data does and doesn't show
  • Investigate how performance and efficiency change as team size, compute and task horizon grow, and find the bottlenecks that limit them
  • Build and scale the systems that run very large agent teams reliably, and debug the failures that only appear at scale
  • Design evaluations for long-horizon problems, and keep their results trustworthy
  • Build the tooling and metrics that let researchers see what a large agent team is doing and why
  • Partner with research teams across Anthropic so they can run their own experiments on the platform, and communicate findings clearly

Technologies

Distributed systemsSchedulersSandboxed code executionInference infrastructureReinforcement Learning infrastructureEvaluationsBenchmarksHarnesses for LLMsLarge language models (LLMs)Scaling laws research

See if your resume is ready for this job

See how our AI can optimize your resume and improve your chances for this role.