Staff+ Software Engineer, Safeguards na Anthropic

Híbrido - San Francisco, CA | New York City, NY

Candidatar-se
Ver mais vagas na Anthropic

Anthropic is seeking a Staff+ Software Engineer for its Safeguards team in San Francisco and New York City. The role focuses on building safety and oversight mechanisms for AI systems, including monitoring models, detecting misuse, and enforcing policies. Engineers will work on agentic systems, intervention tools, and data intelligence, collaborating across teams and infrastructure to protect users and ensure model behavior aligns with safety principles.

Salary

USD 320,000 - 485,000

Requirements

Skills

  • Bachelor’s degree in Computer Science, Software Engineering or comparable experience
  • Proficiency in Python and Typescript
  • Ability to work across the stack
  • Strong communication skills and ability to explain complex technical concepts to non-technical stakeholders
  • 8+ years of experience in a software engineering position
  • Experience with integrity, spam, fraud, or abuse detection and mitigation
  • Experience building trust and safety detection mechanisms and intervention for AI/ML systems
  • Experience with prompt engineering, jailbreak attacks, and other adversarial inputs
  • Worked closely with operational teams to build custom internal tooling

Responsibilities

  • Develop monitoring systems to detect unwanted behaviors from our API partners and potentially take automated enforcement actions; surface these in internal dashboards to analysts for manual review
  • Build abuse detection mechanisms and infrastructure
  • Surface abuse patterns to our research teams to harden models at the training stage
  • Build robust and reliable multi-layered defenses for real-time improvement of safety mechanisms that work at scale

Technologies

PythonTypescript

Compartilhar vaga

Descubra se seu currículo está pronto para esta vaga

Veja como nossa IA pode otimizar seu currículo e aumentar suas chances de conseguir esta posição.