The Cloud Inference team scales and optimizes Claude across major cloud providers. The role focuses on validating inference server and load balancer changes, ensuring performance and reliability, building CI/CD pipelines, and driving model launches across AWS, GCP, and Azure.
Staff + Sr. Software Engineer, Cloud Inference Launch Engineering na Anthropic
Híbrido - San Francisco, CA; Seattle, WA
Ver mais vagas na AnthropicSalary
USD 320,000 - 485,000
Requirements
Skills
- Strong interest in LLM serving; prior inference or ML experience is not required
- Significant software engineering experience with a strong background in high-performance, large-scale distributed systems serving millions of users
- Track record of building automation or test infrastructure that measurably improved release velocity or reliability
- Experience building or operating services on at least one major cloud platform (AWS, GCP, or Azure), with exposure to Kubernetes, Infrastructure as Code, or container orchestration
- Thrive in cross-functional collaboration with both internal teams and external partners
- Fast learner who can quickly ramp up on new technologies, hardware platforms, and provider ecosystems
- Highly autonomous and take ownership of problems end-to-end, including work that falls outside your job description
- LLM inference optimization, batching, and caching strategies
- Capacity-constrained scheduling or shared-resource test infrastructure
- Solid understanding of multi-region deployments, request routing, load balancing, global traffic management
- Working with CSP partner teams to scale infrastructure across multiple platforms, navigating differences in networking, security, privacy, and managed service
- Proficiency in Python or Rust
Responsibilities
- Be on the critical path for frontier model launches, bringing up inference for new model architectures and shipping them to cloud platforms in lockstep with our first‑party platform
- Work with the core inference team to bring new inference features (e.g. structured sampling, prompt caching, and more) to cloud platforms, owning the platform‑specific integration that gets them to production
- Identify and dive deep on the gaps that make inference behave differently across first‑party and CSPs — config drift, observability, deployment patterns, hard cross‑platform bugs — and fix them at the source rather than building platform‑specific workarounds
- Design, build, and own the CI/CD infrastructure for the inference server and load balancer across cloud platforms, with shadow traffic, performance baselines (throughput and latency), and correctness checks that catch regressions before production
- Drive down merge‑to‑production cycle time by making validation faster, more parallel, and cost‑effective enough to run on the same constrained accelerator pool that serves customers, without trading away reliability
- Analyze observability data across providers to identify performance bottlenecks, cost anomalies, and regressions, and drive remediation based on real‑world production workloads
Technologies
PythonRustKubernetesInfrastructure as CodeContainer orchestrationAWSGCPAzureCI/CD infrastructureObservability
Descubra se seu currículo está pronto para esta vaga
Veja como nossa IA pode otimizar seu currículo e aumentar suas chances de conseguir esta posição.