An in-progress recreation of Anthropic’s Towards Monosemanticity, exploring sparse autoencoders and interpretable features in language models.
What I’m Recreating
I’m recreating Towards Monosemanticity: Decomposing Language Models With Dictionary Learning, the 2023 paper by Bricken et al. at Anthropic. The project explores how sparse autoencoders can help us understand what a language model has learned.
I read a really interesting
Current Status
This recreation is in progress!
Reference
Bricken et al. (2023). Towards Monosemanticity: Decomposing Language Models With Dictionary Learning. Transformer Circuits Thread, Anthropic.