Reward model hacking as a challenge for reward learning
📖 Description
This post discusses an issue that could lead to catastrophically misaligned AI even when we have access to a perfect reward signal and there are no misaligned inner optimizers. Instead, the misalignment comes from the fact that our reward signal is too expensive to use directly for RL training, so we train a reward model, which is incorrect on some off-distribution transitions. The agent might then exploit these off-distribution deficiencies, which I'll refer to as *reward model hacking*.
📊 Game Impacts
| Variable | Change | Condition |
|---|---|---|
| Research | +10 | Always |
| Vibey Doom | +2 | Always |
| Ethics Risk | -5 | Always |
💭 Reactions
🔬 Safety Researcher Reaction:
⚠️ Placeholder - Needs Real Quote
"Interesting perspective on safety challenges"
"Interesting perspective on safety challenges"
📰 Media Reaction:
⚠️ Placeholder - Needs Real Quote
"Discussed in AI safety community"
💡 Found a Real Quote? Suggest it here
"Discussed in AI safety community"
🤝 Found an Issue?
This event data is sourced from the pdoom-data repository. If you notice errors or want to suggest improvements:
GitHub Issue (Preferred) 📧 Email (No GitHub)