Research Engineer / Performance Engineer, RL Distributed Systems at Anthropic

On-site - San Francisco, CA; New York City, NY; Seattle, WA

Apply
More jobs at Anthropic

This role is a Research Engineer / Performance Engineer focused on reinforcement learning distributed systems at Anthropic. You will design, build, and operate large-scale distributed systems that support RL training, sampling, and environment execution across diverse hardware, ensuring fault tolerance, autoscaling, and observability while working closely with researchers and performance engineers.

Salary

USD 500,000 - 850,000

Requirements

Skills

  • Strong software engineering skills in Python and at least one systems language such as Rust, C++, or Go
  • Experience designing, building, and operating large-scale distributed systems in production
  • Deep understanding of distributed systems fundamentals, including consistency, coordination, consensus, failure modes, and recovery
  • Ability to reason quantitatively about throughput, latency, and resource costs across compute, memory, storage, and network
  • Experience debugging complex failures across many hosts and services
  • Strong written communication, including design documents and incident writeups

Responsibilities

  • Design, build, and operate the distributed systems that run RL at scale, across training, sampling, and environment execution
  • Find and remove whatever currently limits the system, whether it's scheduling, data movement, storage, networking, or coordination
  • Build fault tolerance into every layer: failure detection, isolation, and recovery that keep long-running jobs making progress without human intervention
  • Design resource management and autoscaling so that compute follows demand as a run's needs shift
  • Build observability that makes it possible to understand what a run is doing and why it slowed down, stalled, or produced unexpected results
  • Build automation that detects and remediates common problems, and design interfaces that let engineers and automated tools operate runs safely
  • Work with researchers and performance engineers to make sure systems changes preserve training correctness and don't introduce subtle nondeterminism
  • Remove classes of failure at their source through incident review, testing, and redesign, and write clear design documents for what you build

Technologies

PythonRustC++GoKubernetesasync Python frameworks (Trio, asyncio)High-performance networkingRDMACollective communication libraries

See if your resume is ready for this job

See how our AI can optimize your resume and improve your chances for this role.