📢

How To Go From Interpretability To Alignment: Just Retarget The Search

📅 2022
public awareness
🔵 Rare
#ai #ai risk #inner alignment #interpretability (ml & ai)

📖 Description

When people talk about [prosaic alignment proposals](https://www.lesswrong.com/posts/fRsjBseRuvRhMPPE5/an-overview-of-11-proposals-for-building-safe-advanced-ai), there's a common pattern: they'll be outlining some overcomplicated scheme, and then they'll say "oh, and assume we have great interpretability tools, this whole thing just works way better the better the interpretability tools are", and then they'll go back to the overcomplicated scheme. (Credit to [Evan](https://www.lesswrong.com/users/evhub) for pointing out this pattern to me.) And then usually there's a whole discussion about the specific problems with the overcomplicated scheme.

📊 Game Impacts

Variable Change Condition
Research +5 Always
Vibey Doom +5 Always
Ethics Risk -5 Always

💭 Reactions

🔬 Safety Researcher Reaction: ⚠️ Placeholder - Needs Real Quote
"Important work advancing our understanding of AI safety"
📰 Media Reaction: ⚠️ Placeholder - Needs Real Quote
"Discussed in AI safety community"
💡 Found a Real Quote? Suggest it here

🔗 Sources

🏷️ Event Metadata

Think this event's metadata could be improved? Suggest changes to category, rarity, tags, game impacts, or p(doom) effects.

🤝 Found an Issue?

This event data is sourced from the pdoom-data repository. If you notice errors or want to suggest improvements:

GitHub Issue (Preferred) 📧 Email (No GitHub)
← Back to All Events