Staff Software Engineer, AI Reliability Engineering en Anthropic

Híbrido - Dublin, Ireland

Postularse
Más vacantes en Anthropic

Join Anthropic’s AI Reliability Engineering team in Dublin as a Staff Software Engineer. You will help keep Claude reliable by developing service level objectives, designing monitoring and observability systems, building high‑availability infrastructure across regions and cloud providers, and leading incident response for critical AI services. This hybrid role blends distributed systems expertise with cross‑team collaboration, requiring strong communication skills, a holistic systems mindset, and a passion for building robust, scalable AI infrastructure.

Salary

EUR 235,000 - 295,000

Requirements

Skills

  • Strong distributed systems, infrastructure, or reliability background
  • Curiosity and bravery to jump into unfamiliar systems during incidents
  • Holistic understanding of how systems compose and where the seams are
  • Ability to build lasting relationships across teams
  • Ownership mindset over outcomes, even for systems you don't own
  • Excellent communication and collaboration skills
  • Diverse experience building product stacks, scaling databases, and running massive distributed systems
  • Bachelor’s degree or equivalent combination of education, training, and experience
  • Experience as an SRE, Production Engineer, or similar reliability-focused role on large scale systems
  • Experience operating large-scale model serving or training infrastructure ( >1000 GPUs )
  • Experience with ML hardware accelerators (GPUs, TPUs, Trainium)
  • Knowledge of ML-specific networking optimizations like RDMA and InfiniBand
  • Expertise in AI-specific observability tools and frameworks
  • Experience with chaos engineering and systematic resilience testing
  • Contribution to open-source infrastructure or ML tooling

Responsibilities

  • Develop appropriate Service Level Objectives for large language model serving systems, balancing availability and latency with development velocity.
  • Design and implement monitoring and observability systems across the token path.
  • Assist in the design and implementation of high-availability serving infrastructure across multiple regions and cloud providers.
  • Lead incident response for critical AI services, ensuring rapid recovery, thorough incident reviews, and systematic improvements.
  • Support the reliability of safeguard model serving – critical for both site reliability and Anthropic's safety commitments.

Technologies

Distributed systemsInfrastructureReliability engineeringMonitoringObservabilityHigh-availability serving infrastructureCloud providersML hardware acceleratorsGPUsTPUsTrainiumRDMAInfiniBandAI-specific observability toolsChaos engineeringOpen-source infrastructureML tooling

Compartir vacante

Descubre si tu currículum está listo para esta vacante

Mira cómo nuestra IA puede optimizar tu currículum y aumentar tus chances en este puesto.