Anthropic’s Infrastructure organization builds and operates the distributed systems that train, serve, and secure AI models. This role works directly with research teams to build reliable, scalable, and performant infrastructure, leading multi-month projects, resolving bottlenecks, and improving operational processes.
Staff+ Software Engineer, Research Systems Engineering en Anthropic
Remoto - Remote-Friendly (Travel-Required)
Más vacantes en AnthropicSalary
USD 320,000 - 485,000
Requirements
Skills
- Experience designing, building, and operating large-scale distributed systems or infrastructure in production
- Experience independently scoping and delivering complex, ambiguous, multi-month technical projects
- Prior experience as a technical lead or mentor for other engineers
- Experience making architectural decisions that other engineers and teams build on top of
- Strong software engineering fundamentals and proficiency in at least one programming language (e.g., Python, Rust, Go, Java)
- Experience with modern cloud infrastructure (e.g., Kubernetes, infrastructure as code, AWS, GCP)
- Strong written and verbal communication skills, with experience building alignment across multiple teams or stakeholders
- 10+ years of software engineering experience, not including internships
- Experience with machine learning infrastructure (e.g., GPUs, TPUs, Trainium) and associated networking infrastructure (e.g., NCCL)
- Low-level systems experience (e.g., Linux kernel tuning, eBPF)
- Experience applying security or privacy engineering best practices
Responsibilities
- Independently scope and lead complex, multi-month infrastructure projects, from an ambiguous starting point through to a production system
- Build deep partnerships with researchers and Research teams to understand their needs and deliver for them
- Mentor other engineers and help raise the technical bar for the team
- Build alignment on technical direction across multiple teams, working through ambiguous problem spaces
- Take ownership of the reliability, scalability, and security of the systems you build as usage and complexity grow
- Lead the improvement of operational processes across Infrastructure, such as incident response, postmortems, and on-call rotations, so the team learns from every incident
Technologies
PythonRustGoJavaKubernetesAWSGCPInfrastructure as CodeNVIDIA GPUsTPUsTrainiumNCCLLinux kernel tuningeBPFSecurity best practicesPrivacy engineering
Descubre si tu currículum está listo para esta vacante
Mira cómo nuestra IA puede optimizar tu currículum y aumentar tus chances en este puesto.