📜

Safety via selection for obedience

📅 2020
policy development
⚪ Common
#ai

📖 Description

[In a previous post](https://www.alignmentforum.org/s/boLPsyNwd6teK5key/p/BXMCgpktdiawT3K5v), I argued that it's plausible that "the most interesting and intelligent behaviour [of AGIs] won't be directly incentivised by their reward functions" - instead, "many of the selection pressures exerted upon them will come from *emergent* interaction dynamics". If I'm right, and the easiest way to build AGI is using [open-ended](https://arxiv.org/abs/2006.07495) environments and reward functions, then we should be less optimistic about using scalable oversight techniques for the purposes of safety - since capabilities researchers won't need good oversight techniques to get to AGI, and most training will occur in environments in which good and bad behaviour aren't well-defined anyway. In this scenario, the best approach to improving safety might involve structural modifications to training environments to change the emergent incentives of agents, as I'll explain in this post.

📊 Game Impacts

Variable Change Condition
Research +10 Always
Vibey Doom +2 Always
Ethics Risk -5 Always

💭 Reactions

🔬 Safety Researcher Reaction: ⚠️ Placeholder - Needs Real Quote
"Useful research for the community"
📰 Media Reaction: ⚠️ Placeholder - Needs Real Quote
"Discussed in AI safety community"
💡 Found a Real Quote? Suggest it here

🔗 Sources

🏷️ Event Metadata

Think this event's metadata could be improved? Suggest changes to category, rarity, tags, game impacts, or p(doom) effects.

🤝 Found an Issue?

This event data is sourced from the pdoom-data repository. If you notice errors or want to suggest improvements:

GitHub Issue (Preferred) 📧 Email (No GitHub)
← Back to All Events