The Safeguards ML Infra team designs, builds, and operates the production infrastructure that powers Claude's safety systems. This role sits at the center of operational work, ensuring safeguards are configured and deployed for model launches, automating runbooks, and building a safeguards registry. The position requires deep experience in production change management at scale and a strong record of shipping safe, critical systems.
Staff+ Site Reliability Engineer, Safeguards ML Infra at Anthropic
Remote - San Francisco, CA; Seattle, WA; New York City, NY; Remote-Friendly
More jobs at AnthropicSalary
USD 405,000 - 485,000
Requirements
Skills
- Production change management at scale
- Deploy pipelines
- Configuration management systems
- Canary analysis
- High‑stakes release experience
- On‑call and incident response
- Postmortem improvement
- AWS cloud experience
- GCP cloud experience
- Python proficiency
- Rust proficiency (plus)
- 8+ years industry software engineering or site reliability engineering experience
- Automation of operational toil
- Launch or production‑readiness review processes across teams
- Familiarity with LLM inference systems and transformer models
Responsibilities
- Launch captain model releases: stand up, configure, and verify safeguards for each new model
- Own off‑cycle deployment of safety classifiers with canary rollouts and post‑deploy validations
- Verify safeguards are live across all platforms (1P, AWS Bedrock, GCP Vertex, etc.) and eliminate configuration drift
- Automate runbooks into tooling and continuous validation pipelines
- Build and maintain a safeguards registry with full provenance
- Participate in on‑call and operational‑duty rotations for incidents and launches
Technologies
PythonRustAWSGCP
See if your resume is ready for this job
See how our AI can optimize your resume and improve your chances for this role.