Recreating Towards Monosemanticity

In Progress Mechanistic Interpretability Sparse Autoencoders

An in-progress recreation of Anthropic’s Towards Monosemanticity, exploring sparse autoencoders and interpretable features in language models.

What I’m Recreating

I’m recreating Towards Monosemanticity: Decomposing Language Models With Dictionary Learning, the 2023 paper by Bricken et al. at Anthropic. The project explores how sparse autoencoders can help us understand what a language model has learned.

I read a really interesting

Current Status

This recreation is in progress!

Reference

Bricken et al. (2023). Towards Monosemanticity: Decomposing Language Models With Dictionary Learning. Transformer Circuits Thread, Anthropic.

Back to Projects