Reverse-engineering using interpretability
📖 Description
Building a model for which you're confident your interpretability is correct, by reverse-engineering each part of the model to work how your interpretability says it should work. (Based on discussion in alignment reading group, ideas from William, Adam, dmz, Evan, Leo, maybe others)
📊 Game Impacts
| Variable | Change | Condition |
|---|---|---|
| Research | +10 | Always |
| Vibey Doom | +2 | Always |
| Ethics Risk | -5 | Always |
💭 Reactions
🔬 Safety Researcher Reaction:
⚠️ Placeholder - Needs Real Quote
"Interesting perspective on safety challenges"
"Interesting perspective on safety challenges"
📰 Media Reaction:
⚠️ Placeholder - Needs Real Quote
"Discussed in AI safety community"
💡 Found a Real Quote? Suggest it here
"Discussed in AI safety community"
🤝 Found an Issue?
This event data is sourced from the pdoom-data repository. If you notice errors or want to suggest improvements:
GitHub Issue (Preferred) 📧 Email (No GitHub)