Staff+ Software Engineer, Research Systems Engineering at Anthropic

Remote - Remote-Friendly (Travel-Required)

Apply
More jobs at Anthropic

Anthropic’s Infrastructure organization builds and operates the distributed systems that train, serve, and secure AI models. This role works directly with research teams to build reliable, scalable, and performant infrastructure, leading multi-month projects, resolving bottlenecks, and improving operational processes.

Salary

USD 320,000 - 485,000

Requirements

Skills

  • Experience designing, building, and operating large-scale distributed systems or infrastructure in production
  • Experience independently scoping and delivering complex, ambiguous, multi-month technical projects
  • Prior experience as a technical lead or mentor for other engineers
  • Experience making architectural decisions that other engineers and teams build on top of
  • Strong software engineering fundamentals and proficiency in at least one programming language (e.g., Python, Rust, Go, Java)
  • Experience with modern cloud infrastructure (e.g., Kubernetes, infrastructure as code, AWS, GCP)
  • Strong written and verbal communication skills, with experience building alignment across multiple teams or stakeholders
  • 10+ years of software engineering experience, not including internships
  • Experience with machine learning infrastructure (e.g., GPUs, TPUs, Trainium) and associated networking infrastructure (e.g., NCCL)
  • Low-level systems experience (e.g., Linux kernel tuning, eBPF)
  • Experience applying security or privacy engineering best practices

Responsibilities

  • Independently scope and lead complex, multi-month infrastructure projects, from an ambiguous starting point through to a production system
  • Build deep partnerships with researchers and Research teams to understand their needs and deliver for them
  • Mentor other engineers and help raise the technical bar for the team
  • Build alignment on technical direction across multiple teams, working through ambiguous problem spaces
  • Take ownership of the reliability, scalability, and security of the systems you build as usage and complexity grow
  • Lead the improvement of operational processes across Infrastructure, such as incident response, postmortems, and on-call rotations, so the team learns from every incident

Technologies

PythonRustGoJavaKubernetesAWSGCPInfrastructure as CodeNVIDIA GPUsTPUsTrainiumNCCLLinux kernel tuningeBPFSecurity best practicesPrivacy engineering

See if your resume is ready for this job

See how our AI can optimize your resume and improve your chances for this role.