Satisficers Tend To Seek Power: Instrumental Convergence Via Retargetability
📖 Description
**Summary**: Why exactly should smart agents tend to usurp their creators? Previous results only apply to optimal agents tending to stay alive and preserve their future options. I extend the power-seeking theorems to apply to many kinds of policy-selection procedures, ranging from planning agents which choose plans with expected utility closest to a randomly generated number, to satisficers, to policies trained by some reinforcement learning algorithms. The key property is not agent optimality--as previously supposed--but is instead the *retargetability of the policy-selection procedure*. These results hint at which kinds of agent cognition and of agent-producing processes are dangerous by default.
📊 Game Impacts
| Variable | Change | Condition |
|---|---|---|
| Research | +5 | Always |
| Vibey Doom | +5 | Always |
| Ethics Risk | -5 | Always |
💭 Reactions
"This is a significant contribution to alignment research"
"Discussed in AI safety community"
🤝 Found an Issue?
This event data is sourced from the pdoom-data repository. If you notice errors or want to suggest improvements:
GitHub Issue (Preferred) 📧 Email (No GitHub)