📜

Outer vs inner misalignment: three framings

📅 2022
policy development
🔵 Rare
#ai #inner alignment #outer alignment

📖 Description

A core concept in the field of AI alignment is a distinction between two types of misalignment: outer misalignment and inner misalignment. Roughly speaking, the outer alignment problem is the problem of specifying an reward function which captures human preferences; and the inner alignment problem is the problem of ensuring that a policy trained on that reward function actually tries to act in accordance with human preferences. (In other words, it's the distinction between aligning the "outer" training signal versus aligning the "inner" policy.) However, the distinction can be difficult to pin down precisely. In this post I'll give three and a half definitions, which each come progressively closer to capturing my current conception of it. I think Framing 1 is a solid starting point; Framings 1.5 and 2 seem like useful refinements, although less concrete; and Framing 3 is fairly speculative. For those who don't already have a solid grasp on the inner-outer misalignment distinction, I...

📊 Game Impacts

Variable Change Condition
Research +10 Always
Vibey Doom +5 Always
Ethics Risk -5 Always

💭 Reactions

🔬 Safety Researcher Reaction: ⚠️ Placeholder - Needs Real Quote
"This is a significant contribution to alignment research"
📰 Media Reaction: ⚠️ Placeholder - Needs Real Quote
"Discussed in AI safety community"
💡 Found a Real Quote? Suggest it here

🔗 Sources

🏷️ Event Metadata

Think this event's metadata could be improved? Suggest changes to category, rarity, tags, game impacts, or p(doom) effects.

🤝 Found an Issue?

This event data is sourced from the pdoom-data repository. If you notice errors or want to suggest improvements:

GitHub Issue (Preferred) 📧 Email (No GitHub)
← Back to All Events