Codebook Features: Sparse and Discrete Interpretability for Neural Networks
📖 Description
CODEBOOK FEATURES : SPARSE AND DISCRETE INTERPRETABILITY FOR NEURAL NETWORKS Alex Tamkin Anthropic?Mohammad Taufeeque FAR AINoah D. Goodman Stanford University ABSTRACT Understanding neural networks is challenging in part because of the dense, con- tinuous nature of their hidden states. We explore whether we can train neural networks to have hidden states that are sparse, discrete, and more interpretable by quantizing their continuous features into what we call codebook features . Code- book features are produced by finetuning neural networks with vector quantization bottlenecks at each layer, producing a network whose hidden features are the sum of a small number of discrete vector codes chosen from a larger codebook. Sur- prisingly, we find that neural networks can operate under this extreme bottleneck with only modest degradation in performance. This sparse, discrete bottleneck also provides an intuitive way of controlling neural network behavior: first, find codes that activate ...
📊 Game Impacts
| Variable | Change | Condition |
|---|---|---|
| Research | +15 | Always |
| Papers | +10 | Always |
| Vibey Doom | +3 | Always |
💭 Reactions
"Notable work on AI safety"
"Published in academic venue"
🔗 Sources
🤝 Found an Issue?
This event data is sourced from the pdoom-data repository. If you notice errors or want to suggest improvements:
GitHub Issue (Preferred) 📧 Email (No GitHub)