Anthropic’s Safeguards organization builds the policies, evaluations, and detection and enforcement systems that define and hold the limits on how Claude can be used. In this role, the Head of Policy Design, Societal Harms will lead a policy design team that manages the consumer harms portfolio—child safety, user well‑being, harmful manipulation, and election integrity. The role involves setting strategy for mitigations built on top of the model, coordinating policy decisions across the portfolio, partnering with cross‑functional teams throughout the model development cycle, engaging external experts and regulators, and serving as the escalation point for high‑severity consumer harms decisions. The position is headquartered in San Francisco and offers competitive compensation, optional equity donation matching, generous vacation and parental leave, flexible working hours, and a collaborative office environment.
Head of Policy Design, Societal Harms en Anthropic
San Francisco, CA
Más vacantes en AnthropicSalary
USD 330,000 - 395,000
Requirements
Skills
- Bachelor’s degree or an equivalent combination of education, training, and/or experience
- Experience leading teams — including managing managers or senior specialists — in AI safety, product policy, or a related field
- Deep, applied familiarity with consumer harm areas such as child safety, mental health and well‑being, manipulation, or election integrity
- Track record of exceptional cross‑team collaboration, building durable working relationships with teams you don’t control
- Working understanding of frontier model development and deployment, including training, fine‑tuning, evaluations, and launch processes
- Experience translating policy positions into enforceable mechanisms and communicating reasoning to technical and non‑technical audiences, including executives
- Sound judgment in ambiguous, high‑consequence decisions, and comfort making a call and escalating appropriately on incomplete information
- Subject‑matter depth in one or more portfolio harm areas from academia, clinical practice, civil society, government, or trust & safety work
- Experience working directly with model training or research teams on model behavior, or shaping the character of a deployed AI system
- Experience with generative AI safety systems, including LLM‑based classification, evaluation, or enforcement pipelines
- Experience engaging external stakeholders in these domains — child safety organizations, election authorities, mental health experts, or regulators
- Experience using agentic AI tools to scale a team’s analysis and operations
Responsibilities
- Lead, develop, and grow the managers and teams responsible for the consumer harms portfolio, including child safety, user well‑being, harmful manipulation, and election integrity
- Coordinate policy decisions across the portfolio, and build the mechanisms that keep them tracked, consistent, and legible
- Set the strategy for how mitigations built on top of the model — policies, detection and enforcement systems, and product interventions — complement what is trained into the model itself
- Prioritize across harm areas competing for the same resources, and make those tradeoffs and their rationale clear to leadership
- Serve as the escalation point for high‑severity and ambiguous consumer harms decisions, including rapid response to emerging risks
- Partner with engineering, data science, product, legal, and research across the model development cycle so consumer harms considerations are represented from training through launch, on every surface where Claude is deployed
- Engage external experts, civil society organizations, and regulators, and translate that engagement into stronger policy and enforcement
Technologies
LLM-based classificationEvaluation pipelinesEnforcement pipelinesAgentic AI toolsModel trainingModel behavior
Descubre si tu currículum está listo para esta vacante
Mira cómo nuestra IA puede optimizar tu currículum y aumentar tus chances en este puesto.