Staff+ Software Engineer, ML Sampling Path na Anthropic

Híbrido - San Francisco, CA

Candidatar-se
Ver mais vagas na Anthropic

The Safeguards ML Sampling Path team builds and operates production services powering Claude's safety systems. You will design, build, and operate backend systems that process tokens on the generation path, manage latency and reliability, and drive performance improvements. The role is hybrid, requiring onsite presence in the San Francisco office at least 25% of the time.

Requirements

Skills

  • Designed, built, and operated high QPS systems at global scale with incident response, outages, and postmortem-driven remediation
  • Strong foundation in distributed systems: replication, consistency tradeoffs, failure modes, and SLO management under load
  • Designed systems for graceful degradation to avoid failures
  • Shipped broad or all-encompassing changes to mission-critical systems (e.g., database migrations, interface changes, rewrites)
  • 8+ years of industry software engineering experience
  • Familiarity with LLM inference systems and transformer-based models (not required, but a plus)
  • Bachelor’s degree or equivalent in a field relevant to the role

Responsibilities

  • Design, build, and operate the backend systems that process every token on the generation path for Claude requests, including the streaming contract with the API and inference engines
  • Own latency and reliability end-to-end: define and maintain SLOs and error budgets for added latency, time-to-first-token, and availability, and lead incident response and postmortem follow-through
  • Ship changes to the hot path rapidly but safely—canaried and gradual rollouts, error budget and latency gating, fast rollbacks—while driving per-token performance
  • Set technical direction for the sampling path: lead design reviews, make latency, reliability, and cost trade-off calls with the inference and research teams, mentor engineers, and raise the operational bar for the wider Safeguards organization

Technologies

Distributed systemsLLM inference systemsTransformer-based modelsAPI streaming contractsSLO managementIncident response and postmortem processes

Compartilhar vaga

Descubra se seu currículo está pronto para esta vaga

Veja como nossa IA pode otimizar seu currículo e aumentar suas chances de conseguir esta posição.