📜

Multi-dimensional rewards for AGI interpretability and control

📅 2021
policy development
⚪ Common
#ai #interpretability (ml & ai) #reinforcement learning

📖 Description

*Update August 2021:* Re-reading this post, I continue to think this is a good and important idea, and I was very happy to learn after I wrote it that what I had in mind here is really a plausible, viable thing to do, even given the cost and performance requirements that people will demand of our future AGIs. I base that belief on the fact that (I now think) the brain does more-or-less exactly what I talk about here (see my post [A model of decision-making in the brain](https://www.lesswrong.com/posts/e5duEqhAhurT8tCyr/a-model-of-decision-making-in-the-brain-the-short-version)), and also on the fact that the machine learning literature also has things like this (see the comments section at the bottom).

📊 Game Impacts

Variable Change Condition
Research +10 Always
Ethics Risk -5 Always

💭 Reactions

🔬 Safety Researcher Reaction: ⚠️ Placeholder - Needs Real Quote
"Background research"
📰 Media Reaction: ⚠️ Placeholder - Needs Real Quote
"Discussed in AI safety community"
💡 Found a Real Quote? Suggest it here

🔗 Sources

🏷️ Event Metadata

Think this event's metadata could be improved? Suggest changes to category, rarity, tags, game impacts, or p(doom) effects.

🤝 Found an Issue?

This event data is sourced from the pdoom-data repository. If you notice errors or want to suggest improvements:

GitHub Issue (Preferred) 📧 Email (No GitHub)
← Back to All Events