Instrumental Convergence For Realistic Agent Objectives
📖 Description
**Edit, 5/16/23: I think this post is beautiful, correct in its narrow technical claims, and practically irrelevant to alignment. This post presents an unrealistic picture of the role of reward functions in reinforcement learning, conflating "utility" with "reward." Reward functions are not "goals", reward functions are not "objectives" of the policy network, real-world policies are not "optimal", and the mechanistic function of reward is (usually) to provide policy gradients to update the policy network.**
📊 Game Impacts
| Variable | Change | Condition |
|---|---|---|
| Research | +5 | Always |
| Vibey Doom | +2 | Always |
| Ethics Risk | -5 | Always |
💭 Reactions
🔬 Safety Researcher Reaction:
⚠️ Placeholder - Needs Real Quote
"Adds to our knowledge base"
"Adds to our knowledge base"
📰 Media Reaction:
⚠️ Placeholder - Needs Real Quote
"Discussed in AI safety community"
💡 Found a Real Quote? Suggest it here
"Discussed in AI safety community"
🤝 Found an Issue?
This event data is sourced from the pdoom-data repository. If you notice errors or want to suggest improvements:
GitHub Issue (Preferred) 📧 Email (No GitHub)