The Cloud Inference team scales and optimizes Claude to serve developers and enterprise customers across multiple cloud providers. Engineers build end-to-end services from API integration to inference execution, managing compute, networking, and operational models. They design backend services, collaborate cross-functionally, build CI/CD pipelines, create cost-effective tooling, and contribute to capacity planning, autoscaling, and routing strategies to maintain performance and safety at massive scale.
Staff + Sr. Software Engineer, Cloud Inference na Anthropic
Híbrido - San Francisco, CA, USA
Ver mais vagas na AnthropicSalary
USD 320,000 - 485,000
Requirements
Skills
- Significant software engineering experience in high-performance, large-scale distributed systems serving millions of users
- Experience building or operating services on at least one major cloud platform (AWS, GCP, or Azure) with exposure to Kubernetes, Infrastructure as Code, or container orchestration
- Curiosity about LLM serving; prior inference or ML experience not required
- Strong cross-functional collaboration with internal and external teams
- Experience working with external partners to align goals and deliver impact
- Fast learner who can quickly ramp up on new technologies, hardware platforms, and provider ecosystems
- High autonomy and ownership of end-to-end problem solving
- Direct experience working with CSPs to scale infrastructure or products across multiple platforms
- Hands‑on experience with capacity management, cost optimization, or resource planning at scale across heterogeneous environments
- Solid understanding of multi-region deployments, geographic routing, and global traffic management
- Proficiency in Python or Rust
- Bachelor’s degree or equivalent education, training, and/or experience
Responsibilities
- Design, build, and own backend services and infrastructure that serve Claude across multiple CSPs, accounting for differences in compute hardware, networking, APIs, and operational models
- Work cross-functionally with internal inference, product API, systems, and security teams, and with CSP partners to stand up the full serving stack on new cloud platforms, resolve operational issues, and influence provider roadmaps
- Build and evolve CI/CD automation systems, including validation and deployment pipelines, that reliably ship new model versions to millions of users across cloud platforms without regressions
- Design interfaces and tooling abstractions across CSPs that enable cost-effective inference management, scale across providers, and reduce per-platform complexity
- Contribute to capacity planning, autoscaling, and workload routing strategies that match supply with demand and direct requests to the most cost-effective accelerator and region
- Analyze observability data across providers to identify performance bottlenecks, cost anomalies, and regressions, and drive remediation based on real-world production workloads
Technologies
PythonRustKubernetesInfrastructure as CodeContainer orchestrationCI/CD automationCapacity planningAutoscalingWorkload routingObservability tools
Descubra se seu currículo está pronto para esta vaga
Veja como nossa IA pode otimizar seu currículo e aumentar suas chances de conseguir esta posição.