Staff Software Engineer, AI Reliability en Anthropic

Híbrido - San Francisco, CA | New York City, NY | Seattle, WA

Postularse
Más vacantes en Anthropic

Anthropic’s AI Reliability Engineering (AIRE) team focuses on enhancing the reliability of critical serving paths for Claude, including SDKs, API layers, infrastructure, and accelerators. The Staff Software Engineer, AI Reliability will partner with cross‑functional teams to design Service Level Objectives, build monitoring and observability solutions, develop high‑availability infrastructure across regions, lead incident response for key AI services, and ensure the robustness of safeguard model serving. This role demands a deep understanding of distributed systems, reliability engineering, and AI infrastructure, with opportunities to impact large‑scale model serving and training environments.

Salary

USD 325,000 - 485,000

Requirements

Skills

  • Strong distributed systems, infrastructure, or reliability background
  • Curiosity and bravery to tackle unfamiliar systems during incidents
  • Holistic thinking about how systems compose and where seams are
  • Ability to build lasting relationships across teams
  • Ownership over outcomes, even for systems you don't own
  • Excellent communication and collaboration skills
  • Diverse experience across product stacks, scaled databases, and large distributed systems
  • SRE, Production Engineer, or similar reliability-focused role experience
  • Experience operating large-scale model serving or training infrastructure (>1000 GPUs)
  • Experience with ML hardware accelerators (GPUs, TPUs, Trainium)
  • Understanding of ML-specific networking optimizations like RDMA and InfiniBand
  • Expertise in AI-specific observability tools and frameworks
  • Experience with chaos engineering and systematic resilience testing
  • Contributions to open-source infrastructure or ML tooling
  • Minimum education: Bachelor’s degree or equivalent
  • Required field of study: relevant to the role

Responsibilities

  • Develop appropriate Service Level Objectives for large language model serving systems, balancing availability and latency with development velocity
  • Design and implement monitoring and observability systems across the token path
  • Assist in the design and implementation of high-availability serving infrastructure across multiple regions and cloud provider
  • Lead incident response for critical AI services, ensuring rapid recovery, thorough incident reviews, and systematic improvements
  • Support the reliability of safeguard model serving – critical for both site reliability and Anthropic's safety commitments

Technologies

GPUsTPUsTrainiumRDMAInfiniBandAI-specific observability toolsChaos engineeringOpen-source infrastructure or ML tooling

Compartir vacante

Descubre si tu currículum está listo para esta vacante

Mira cómo nuestra IA puede optimizar tu currículum y aumentar tus chances en este puesto.