📜

Reward model hacking as a challenge for reward learning

📅 2022
policy development
⚪ Common
#ai #reward functions

📖 Description

This post discusses an issue that could lead to catastrophically misaligned AI even when we have access to a perfect reward signal and there are no misaligned inner optimizers. Instead, the misalignment comes from the fact that our reward signal is too expensive to use directly for RL training, so we train a reward model, which is incorrect on some off-distribution transitions. The agent might then exploit these off-distribution deficiencies, which I'll refer to as *reward model hacking*.

📊 Game Impacts

Variable Change Condition
Research +10 Always
Vibey Doom +2 Always
Ethics Risk -5 Always

💭 Reactions

🔬 Safety Researcher Reaction: ⚠️ Placeholder - Needs Real Quote
"Interesting perspective on safety challenges"
📰 Media Reaction: ⚠️ Placeholder - Needs Real Quote
"Discussed in AI safety community"
💡 Found a Real Quote? Suggest it here

🔗 Sources

🏷️ Event Metadata

Think this event's metadata could be improved? Suggest changes to category, rarity, tags, game impacts, or p(doom) effects.

🤝 Found an Issue?

This event data is sourced from the pdoom-data repository. If you notice errors or want to suggest improvements:

GitHub Issue (Preferred) 📧 Email (No GitHub)
← Back to All Events