The Staff Software Engineer, Inference role at Anthropic focuses on building and maintaining the critical inference systems that serve Claude to millions of users worldwide. Responsibilities include designing intelligent routing algorithms, autoscaling compute fleets, building production-grade deployment pipelines, integrating new AI accelerator platforms, and supporting inference for new model architectures while optimizing performance and enabling breakthrough research.
Staff Software Engineer, Inference at Anthropic
Hybrid - Dublin, IE
More jobs at AnthropicSalary
EUR 295,000 - 355,000
Requirements
Skills
- High-performance, large-scale distributed systems
- Implementing and deploying machine learning systems at scale
- Load balancing, request routing, or traffic management systems
- LLM inference optimization, batching, and caching strategies
- Kubernetes and cloud infrastructure (AWS, GCP)
- Python or Rust
- Significant software engineering experience, particularly with distributed systems
- Bachelor’s degree or an equivalent combination of education, training, and/or experience
Responsibilities
- Build and maintain critical systems that serve Claude to millions of users worldwide
- Serve models via compute-agnostic inference deployments
- Design intelligent routing algorithms that optimize request distribution across thousands of accelerators
- Autoscale compute fleet to match supply with demand across production, research, and experimental workloads
- Build production-grade deployment pipelines for releasing new models to millions of users
- Integrate new AI accelerator platforms to maintain hardware-agnostic advantage
- Contribute to new inference features such as structured sampling and prompt caching
- Support inference for new model architectures
- Analyze observability data to tune performance based on real-world production workloads
- Manage multi-region deployments and geographic routing for global customers
- Maximize compute efficiency while enabling breakthrough AI research
Technologies
PythonRustKubernetesAWSGCPLLM inferenceBatching strategiesCaching strategiesDistributed systems
See if your resume is ready for this job
See how our AI can optimize your resume and improve your chances for this role.