Anthropic is looking for a Staff Engineer to serve as the technical lead for its Inference Runtime team. The role owns the architecture, roadmap, and performance of the accelerator‑agnostic runtime that powers Claude’s inference serving stack across GPUs, TPUs, and Trainium. The position requires deep systems and ML infrastructure expertise, hands‑on work in Rust and Python, and a strong background in performance optimization, large‑scale distributed systems, and accelerator ecosystems. The engineer will partner with cross‑functional teams, drive technical direction, mentor peers, and ensure the runtime remains efficient, scalable, and platform‑agnostic.
Staff+ Software Engineer, Inference Runtime na Anthropic
Híbrido - Remote-Friendly (Travel-Required), San Francisco, CA, Seattle, WA, New York City, NY
Ver mais vagas na AnthropicSalary
USD 405,000 - 485,000
Requirements
Skills
- Deep background in systems engineering or ML infrastructure, with hands‑on performance profiling, latency and throughput optimization, and large‑scale systems debugging
- Real depth in at least one accelerator ecosystem (CUDA/GPU, TPU, or Trainium/AWS Neuron) and appetite for keeping the runtime agnostic across all of them
- Significant software engineering experience in high‑performance, large‑scale distributed systems serving millions of users
- Track record of defining and using engineering metrics to drive improvement (SLOs, escape rates, release times, latency, throughput)
- Experience driving technical alignment across organizational boundaries and advocating for team needs while contributing to shared infrastructure
- Strong written and verbal communication, and ability to influence technical direction without formal authority
- 8+ years of software engineering experience, including time as a technical lead or anchor on a platform, inference runtime, or ML infrastructure team
- Experience with ML compiler toolchains (XLA, Triton, NeuronX) or accelerator driver/firmware management at scale
- Background operating a production validation surface at scale (shadow traffic, canary populations, automated baseline comparison, fast rollback)
- Experience with deterministic or simulation‑based testing for hardware‑dependent systems
- Experience with CI/CD systems at scale, particularly for workloads involving accelerator hardware
- Familiarity with Kubernetes‑based development and job scheduling environments
Responsibilities
- Set technical direction for the team, owning the architecture and roadmap for the shared runtime of the inference serving stack
- Own and evolve the accelerator‑agnostic runtime itself – its interfaces, internal boundaries, and build structure – including hands‑on work in a performance‑sensitive Rust and Python codebase
- Keep the platform’s expansion cost low by ensuring new models and deployment targets pay only for their own specialization, and edge cases stitch back into the core easily
- Drive efficient accelerator usage – utilization, scheduling, memory management – across GPU, TPU, and Trainium
- Build the runtime’s validation surface around partitioned builds, change‑scoped testing, and canary/shadow/rollback as first‑class mechanisms
- Act as a technical counterpart to Anthropic’s central Infrastructure org on the compilers, build systems, and toolchains the runtime depends on, contributing Inference’s performance and correctness requirements, and making the call on build vs. adopt
- Mentor engineers on the team through design review, code review, and direct collaboration, raising the technical bar without owning headcount
Technologies
RustPythonCUDAGPUTPUTrainiumXLATritonNeuronXKubernetesCI/CDML compiler toolchainsBuild systemsToolchainsValidation surface
Descubra se seu currículo está pronto para esta vaga
Veja como nossa IA pode otimizar seu currículo e aumentar suas chances de conseguir esta posição.