📜

How an alien theory of mind might be unlearnable

📅 2022
policy development
⚪ Common
#ai #value learning

📖 Description

**EDIT**: *This is a post about an alien mind being unlearnable in practice. As a reminder, theory of mind is unlearnable in theory, as stated [here](https://arxiv.org/abs/1712.05812) - there is more information in "preferences + (ir)rationality" than there is in "behaviour", "policy", or even "[complete internal brain structure](https://www.lesswrong.com/posts/9rjW9rhyhJijHTM92/learning-human-preferences-black-box-white-box-and)". This information gap must be covered by assumptions (or "labelled data", in CS terms) of one form or another - assumptions that cannot be deduced from observation. It is unclear whether we need only a few trivial assumptions or a lot of detailed and subtle ones. Hence posts like this one, looking at the practicality angle.*

📊 Game Impacts

Variable Change Condition
Research +10 Always
Vibey Doom +2 Always
Ethics Risk -5 Always

💭 Reactions

🔬 Safety Researcher Reaction: ⚠️ Placeholder - Needs Real Quote
"Useful research for the community"
📰 Media Reaction: ⚠️ Placeholder - Needs Real Quote
"Discussed in AI safety community"
💡 Found a Real Quote? Suggest it here

🔗 Sources

🏷️ Event Metadata

Think this event's metadata could be improved? Suggest changes to category, rarity, tags, game impacts, or p(doom) effects.

🤝 Found an Issue?

This event data is sourced from the pdoom-data repository. If you notice errors or want to suggest improvements:

GitHub Issue (Preferred) 📧 Email (No GitHub)
← Back to All Events