This role is part of Anthropic’s Software Engineering – Infrastructure team focused on storage. The position involves designing, operating, and scaling blob storage platforms at hyper‑scale, leading architectural decisions, collaborating across research, training, infrastructure, business, and data teams, and ensuring secure, performant, and cost‑effective storage solutions to support AI workloads.
Staff+ Software Engineer, Storage + Transfer at Anthropic
Hybrid - San Francisco, CA, USA
More jobs at AnthropicSalary
USD 1 - 2
Requirements
Skills
- Experience designing, building, and operating a large-scale distributed storage system in production
- Deep understanding of storage system fundamentals: replication, erasure coding, consistency models, metadata management, durability, failure handling, high availability, access controls, identity models, security approaches, abstraction layers, network requirements, multi-region considerations
- Experience serving as an owner, tech lead, or architect for a large, complex infrastructure system
- Strong software engineering fundamentals and hands‑on coding ability in at least one systems language (C++, Rust, Go, or Java)
- Experience operating and steering critical infrastructure at scale, including security, on‑call, incident response, capacity planning, customer support, and cost efficiency
- Strong written and verbal communication skills, with experience driving alignment on technical direction and user-facing aspects across multiple orgs and stakeholders
Responsibilities
- Shape the technical strategy and architecture for Anthropic's storage layers
- Build strong relationships with users and partner teams to understand their access patterns and unique security, usability, and business requirements
- Partner closely with CSPs, networking, and datacenter teams on infrastructure primitives
- Translate user needs into scalable, achievable system designs and drive alignment
- Make principled tradeoffs across durability, availability, consistency, performance, security, and cost, and document the reasoning
- Break large problems into deliverable milestones, and lead cross‑functional teams to ship new capabilities safely at unprecedented speed
- Work across backend stacks, abstraction layers, and clients to provide users with simple, consistent interfaces no matter where data lives
- Plan and lead large migrations of critical workloads with minimal user disruption
- Participate in and improve operations, including SLOs, observability, capacity planning, incident response, and on‑call
- Stay hands‑on in code and production, including in the most critical and complex areas
Technologies
C++RustGoJava
See if your resume is ready for this job
See how our AI can optimize your resume and improve your chances for this role.