Staff+ Software Engineer in Anthropic’s Safeguards ML Infra team based in San Francisco. The role involves designing, building, and operating production infrastructure that powers Claude’s safety systems across multiple deployment platforms. Responsibilities include developing backend services, maintaining SLOs, building observability, leading incident response, automating operations, and collaborating with ML researchers to productionize safety techniques. Candidates should have strong distributed systems background, experience with high‑QPS services, on‑call ownership, and cloud platform operations on AWS and GCP. The team values platform‑agnostic tooling, automation, and reducing operational toil.
Staff+ Software Engineer, Safeguards ML Infrastructure at Anthropic
Hybrid - San Francisco, CA
More jobs at AnthropicSalary
USD 320,000 - 485,000
Requirements
Skills
- Proficient in Python
- Experience with Rust is a plus
- Designed, built, and operated high QPS systems at global scale
- Strong foundation in distributed systems: replication, consistency tradeoffs, failure modes, and SLO management under load
- Meaningful on-call experience for production systems, including incident response and postmortem-driven improvements
- Hands‑on experience deploying and operating on cloud platforms (AWS, GCP) at scale
- Approach infrastructure as a platform – building systems and abstractions for other engineers
- 8+ years of industry software engineering experience
- Experience building deployment and rollout systems with canary analysis, automated validation, or progressive rollout controls
- History of reducing operational toil through automation, including transitioning teams from manual deployment processes to self‑serve pipelines
- Familiarity with LLM inference systems and the operational characteristics of transformer‑based models
Responsibilities
- Design, build, and deploy backend services that are critical safety pieces on the token sampling and generation path
- Own and operate the production serving infrastructure for those services across multiple deployment platforms (1P, AWS Bedrock, GCP Vertex)
- Define and maintain SLOs, build observability and alerting systems, and lead incident response for infrastructure on the critical path of every Claude request
- Participate in on‑call and operational‑duty rotations covering service incidents, model provisioning, and time‑sensitive research and safety launches
- Reduce on‑call and on‑duty toil by building automation, tooling, and self‑serve workflows that minimize manual operations
- Build and maintain a safety registry with full provenance – tracking what is running in production, on which model, and when and by whom it was deployed
- Implement automated post‑deploy validation to ensure correctness is consistent across platforms
- Work closely with ML researchers to productionize new safety techniques, translating experimental work into reliable, scalable production systems
- Contribute to the long‑term goal of platform‑agnostic deployment tooling that brings 3P platforms to parity with 1P operational maturity
Technologies
PythonRustAWSGCP1PBedrockVertexDistributed systemsObservabilityIncident responseAutomation toolingLLM inference systems
See if your resume is ready for this job
See how our AI can optimize your resume and improve your chances for this role.