A Longlist of Theories of Impact for Interpretability
📖 Description
I hear a lot of different arguments floating around for exactly how mechanistically interpretability research will reduce x-risk. As an interpretability researcher, forming clearer thoughts on this is pretty important to me! As a preliminary step, I've compiled a list with a longlist of 19 different arguments I've heard for why interpretability matters. These are pretty scattered and early stage thoughts (and emphatically my personal opinion than the official opinion of Anthropic!), but I'm sharing them in the hopes that this is interesting to people
📊 Game Impacts
| Variable | Change | Condition |
|---|---|---|
| Research | +5 | Always |
| Vibey Doom | +5 | Always |
| Ethics Risk | -5 | Always |
💭 Reactions
🔬 Safety Researcher Reaction:
⚠️ Placeholder - Needs Real Quote
"Important work advancing our understanding of AI safety"
"Important work advancing our understanding of AI safety"
📰 Media Reaction:
⚠️ Placeholder - Needs Real Quote
"Discussed in AI safety community"
💡 Found a Real Quote? Suggest it here
"Discussed in AI safety community"
🤝 Found an Issue?
This event data is sourced from the pdoom-data repository. If you notice errors or want to suggest improvements:
GitHub Issue (Preferred) 📧 Email (No GitHub)