Transformer Circuits
📖 Description
[Chris Olah](https://www.alignmentforum.org/users/christopher-olah), [Neel Nanda](https://www.alignmentforum.org/users/neel-nanda-1), [Catherine Olsson](https://www.lesswrong.com/users/catherio), Nelson Elhage, and a bunch of other people at [Anthropic](https://www.anthropic.com/) just published "Transformer Circuits," an application of the [*Circuits*-style](https://www.alignmentforum.org/posts/MG4ZjWQDrdpgeu8wG/zoom-in-an-introduction-to-circuits) interpretability paradigm to transformer-based language models. From their very top level summary:
📊 Game Impacts
| Variable | Change | Condition |
|---|---|---|
| Research | +5 | Always |
| Vibey Doom | +2 | Always |
| Ethics Risk | -5 | Always |
💭 Reactions
🔬 Safety Researcher Reaction:
⚠️ Placeholder - Needs Real Quote
"Useful research for the community"
"Useful research for the community"
📰 Media Reaction:
⚠️ Placeholder - Needs Real Quote
"Discussed in AI safety community"
💡 Found a Real Quote? Suggest it here
"Discussed in AI safety community"
🤝 Found an Issue?
This event data is sourced from the pdoom-data repository. If you notice errors or want to suggest improvements:
GitHub Issue (Preferred) 📧 Email (No GitHub)