🔬

Codebook Features: Sparse and Discrete Interpretability for Neural Networks

📅 2023
technical research breakthrough
🔵 Rare

📖 Description

CODEBOOK FEATURES : SPARSE AND DISCRETE INTERPRETABILITY FOR NEURAL NETWORKS Alex Tamkin Anthropic?Mohammad Taufeeque FAR AINoah D. Goodman Stanford University ABSTRACT Understanding neural networks is challenging in part because of the dense, con- tinuous nature of their hidden states. We explore whether we can train neural networks to have hidden states that are sparse, discrete, and more interpretable by quantizing their continuous features into what we call codebook features . Code- book features are produced by finetuning neural networks with vector quantization bottlenecks at each layer, producing a network whose hidden features are the sum of a small number of discrete vector codes chosen from a larger codebook. Sur- prisingly, we find that neural networks can operate under this extreme bottleneck with only modest degradation in performance. This sparse, discrete bottleneck also provides an intuitive way of controlling neural network behavior: first, find codes that activate ...

📊 Game Impacts

Variable Change Condition
Research +15 Always
Papers +10 Always
Vibey Doom +3 Always

💭 Reactions

🔬 Safety Researcher Reaction: ⚠️ Placeholder - Needs Real Quote
"Notable work on AI safety"
📰 Media Reaction: ⚠️ Placeholder - Needs Real Quote
"Published in academic venue"
💡 Found a Real Quote? Suggest it here

🔗 Sources

🏷️ Event Metadata

Think this event's metadata could be improved? Suggest changes to category, rarity, tags, game impacts, or p(doom) effects.

🤝 Found an Issue?

This event data is sourced from the pdoom-data repository. If you notice errors or want to suggest improvements:

GitHub Issue (Preferred) 📧 Email (No GitHub)
← Back to All Events