The Interpretability team at Anthropic is focused on reverse‑engineering how trained language models operate, aiming to develop a mechanistic understanding of neural networks. The role seeks researchers and engineers who can devise methods to dissect large language models, design and execute experiments at both toy and large‑scale levels, create and analyze interpretability features and circuits, build supporting infrastructure, and effectively communicate findings internally and to the broader scientific community.
Research Scientist, Interpretability na Anthropic
Presencial - San Francisco, CA
Ver mais vagas na AnthropicResponsibilities
- Develop methods for understanding LLMs by reverse engineering algorithms learned in their weights
- Design and run robust experiments, both quickly in toy scenarios and at scale in large models
- Create and analyze new interpretability features and circuits to better understand how models work
- Build infrastructure for running experiments and visualizing results
- Work with colleagues to communicate results internally and publicly
Technologies
Transformer circuitsNeural networksHaikuSonnet
Descubra se seu currículo está pronto para esta vaga
Veja como nossa IA pode otimizar seu currículo e aumentar suas chances de conseguir esta posição.