Multi-dimensional rewards for AGI interpretability and control
📖 Description
*Update August 2021:* Re-reading this post, I continue to think this is a good and important idea, and I was very happy to learn after I wrote it that what I had in mind here is really a plausible, viable thing to do, even given the cost and performance requirements that people will demand of our future AGIs. I base that belief on the fact that (I now think) the brain does more-or-less exactly what I talk about here (see my post [A model of decision-making in the brain](https://www.lesswrong.com/posts/e5duEqhAhurT8tCyr/a-model-of-decision-making-in-the-brain-the-short-version)), and also on the fact that the machine learning literature also has things like this (see the comments section at the bottom).
📊 Game Impacts
| Variable | Change | Condition |
|---|---|---|
| Research | +10 | Always |
| Ethics Risk | -5 | Always |
💭 Reactions
"Background research"
"Discussed in AI safety community"
🤝 Found an Issue?
This event data is sourced from the pdoom-data repository. If you notice errors or want to suggest improvements:
GitHub Issue (Preferred) 📧 Email (No GitHub)