Anthropic’s AI Reliability Engineering (AIRE) team focuses on enhancing the reliability of critical serving paths for Claude, including SDKs, API layers, infrastructure, and accelerators. The Staff Software Engineer, AI Reliability will partner with cross‑functional teams to design Service Level Objectives, build monitoring and observability solutions, develop high‑availability infrastructure across regions, lead incident response for key AI services, and ensure the robustness of safeguard model serving. This role demands a deep understanding of distributed systems, reliability engineering, and AI infrastructure, with opportunities to impact large‑scale model serving and training environments.
Staff Software Engineer, AI Reliability at Anthropic
Hybrid - San Francisco, CA | New York City, NY | Seattle, WA
More jobs at AnthropicSalary
USD 325,000 - 485,000
Requirements
Skills
- Strong distributed systems, infrastructure, or reliability background
- Curiosity and bravery to tackle unfamiliar systems during incidents
- Holistic thinking about how systems compose and where seams are
- Ability to build lasting relationships across teams
- Ownership over outcomes, even for systems you don't own
- Excellent communication and collaboration skills
- Diverse experience across product stacks, scaled databases, and large distributed systems
- SRE, Production Engineer, or similar reliability-focused role experience
- Experience operating large-scale model serving or training infrastructure (>1000 GPUs)
- Experience with ML hardware accelerators (GPUs, TPUs, Trainium)
- Understanding of ML-specific networking optimizations like RDMA and InfiniBand
- Expertise in AI-specific observability tools and frameworks
- Experience with chaos engineering and systematic resilience testing
- Contributions to open-source infrastructure or ML tooling
- Minimum education: Bachelor’s degree or equivalent
- Required field of study: relevant to the role
Responsibilities
- Develop appropriate Service Level Objectives for large language model serving systems, balancing availability and latency with development velocity
- Design and implement monitoring and observability systems across the token path
- Assist in the design and implementation of high-availability serving infrastructure across multiple regions and cloud provider
- Lead incident response for critical AI services, ensuring rapid recovery, thorough incident reviews, and systematic improvements
- Support the reliability of safeguard model serving – critical for both site reliability and Anthropic's safety commitments
Technologies
GPUsTPUsTrainiumRDMAInfiniBandAI-specific observability toolsChaos engineeringOpen-source infrastructure or ML tooling
See if your resume is ready for this job
See how our AI can optimize your resume and improve your chances for this role.