Anthropic is building a high‑performance inference engine for its Claude AI system. The Performance Engineer will build and optimize the software that routes requests to accelerators, manages memory, and coordinates between host and device. The role focuses on throughput, cost, reliability, and latency across accelerator and cloud platforms, requiring deep understanding of LLM inference, accelerator programming, and high‑performance systems.
Performance Engineer, Inference Engine at Anthropic
Hybrid - San Francisco, CA, United States
More jobs at AnthropicSalary
USD 350,000 - 850,000
Requirements
Skills
- Working mental model of LLM inference: how prefill and decode land on an accelerator’s compute, memory, and interconnect, and what the host is doing meanwhile
- Proven quick learner: ramped fast in deep, unfamiliar systems and shipped consequential changes quickly
- Strong systems programming (Rust, C++, or similar), with care for code quality and tests
- Analytical about performance: observe and profile first, form a hypothesis, test it, then change the code and measure again
- Low ego: ask the naive question, take feedback well, pick up slack outside your job description
- Enjoy pair programming (we love to pair!) and care about the societal impacts of your work
- Bachelor’s degree or an equivalent combination of education, training, and/or experience
Responsibilities
- Build and optimize the inference engine at Anthropic scale, improving throughput, cost, reliability, and latency across accelerator and cloud platforms
- Model performance and make data‑driven changes to the system
- Keep device utilization high by minimizing overhead and ensuring accelerators are fully utilized
- Reuse model state to avoid recomputation and reduce latency
- Build observability to identify performance gaps and drive iterative improvements
- Ensure token quality and safety, working closely with safeguards and safety teams
- Coordinate between host and device in high‑performance, distributed systems
Technologies
RustC++GPU programmingAccelerator programmingOS internalsTransformer architectureAllocatorCacheSchedulerHigh‑bandwidth transport
See if your resume is ready for this job
See how our AI can optimize your resume and improve your chances for this role.