This role is a Research Engineer / Performance Engineer focused on reinforcement learning distributed systems at Anthropic. You will design, build, and operate large-scale distributed systems that support RL training, sampling, and environment execution across diverse hardware, ensuring fault tolerance, autoscaling, and observability while working closely with researchers and performance engineers.
Research Engineer / Performance Engineer, RL Distributed Systems na Anthropic
Presencial - San Francisco, CA; New York City, NY; Seattle, WA
Ver mais vagas na AnthropicSalary
USD 500,000 - 850,000
Requirements
Skills
- Strong software engineering skills in Python and at least one systems language such as Rust, C++, or Go
- Experience designing, building, and operating large-scale distributed systems in production
- Deep understanding of distributed systems fundamentals, including consistency, coordination, consensus, failure modes, and recovery
- Ability to reason quantitatively about throughput, latency, and resource costs across compute, memory, storage, and network
- Experience debugging complex failures across many hosts and services
- Strong written communication, including design documents and incident writeups
Responsibilities
- Design, build, and operate the distributed systems that run RL at scale, across training, sampling, and environment execution
- Find and remove whatever currently limits the system, whether it's scheduling, data movement, storage, networking, or coordination
- Build fault tolerance into every layer: failure detection, isolation, and recovery that keep long-running jobs making progress without human intervention
- Design resource management and autoscaling so that compute follows demand as a run's needs shift
- Build observability that makes it possible to understand what a run is doing and why it slowed down, stalled, or produced unexpected results
- Build automation that detects and remediates common problems, and design interfaces that let engineers and automated tools operate runs safely
- Work with researchers and performance engineers to make sure systems changes preserve training correctness and don't introduce subtle nondeterminism
- Remove classes of failure at their source through incident review, testing, and redesign, and write clear design documents for what you build
Technologies
PythonRustC++GoKubernetesasync Python frameworks (Trio, asyncio)High-performance networkingRDMACollective communication libraries
Descubra se seu currículo está pronto para esta vaga
Veja como nossa IA pode otimizar seu currículo e aumentar suas chances de conseguir esta posição.