How an alien theory of mind might be unlearnable
📖 Description
**EDIT**: *This is a post about an alien mind being unlearnable in practice. As a reminder, theory of mind is unlearnable in theory, as stated [here](https://arxiv.org/abs/1712.05812) - there is more information in "preferences + (ir)rationality" than there is in "behaviour", "policy", or even "[complete internal brain structure](https://www.lesswrong.com/posts/9rjW9rhyhJijHTM92/learning-human-preferences-black-box-white-box-and)". This information gap must be covered by assumptions (or "labelled data", in CS terms) of one form or another - assumptions that cannot be deduced from observation. It is unclear whether we need only a few trivial assumptions or a lot of detailed and subtle ones. Hence posts like this one, looking at the practicality angle.*
📊 Game Impacts
| Variable | Change | Condition |
|---|---|---|
| Research | +10 | Always |
| Vibey Doom | +2 | Always |
| Ethics Risk | -5 | Always |
💭 Reactions
"Useful research for the community"
"Discussed in AI safety community"
🤝 Found an Issue?
This event data is sourced from the pdoom-data repository. If you notice errors or want to suggest improvements:
GitHub Issue (Preferred) 📧 Email (No GitHub)