📢

Benchmark for successful concept extrapolation/avoiding goal misgeneralization

📅 2022
public awareness
⚪ Common
#ai

📖 Description

If an AI has been trained on data about adults, what should it do when it encounters a child? When an AI encounters a situation that its training data hasn't covered, there's a risk that it will incorrectly generalize from what it has been trained for and do the wrong thing. [Aligned AI](https://buildaligned.ai/news/aligned-ai-releases-new-disambiguation-benchmark/) has released a new [benchmark](https://github.com/alignedai/HappyFaces) designed to measure how well image-classifying algorithms avoid goal misgeneralization. This post explains what goal misgeneralization is, what the new benchmark measures, and how the benchmark is related to goal misgeneralization and concept extrapolation.

📊 Game Impacts

Variable Change Condition
Research +5 Always

💭 Reactions

🔬 Safety Researcher Reaction: ⚠️ Placeholder - Needs Real Quote
"Background research"
📰 Media Reaction: ⚠️ Placeholder - Needs Real Quote
"Discussed in AI safety community"
💡 Found a Real Quote? Suggest it here

🔗 Sources

🏷️ Event Metadata

Think this event's metadata could be improved? Suggest changes to category, rarity, tags, game impacts, or p(doom) effects.

🤝 Found an Issue?

This event data is sourced from the pdoom-data repository. If you notice errors or want to suggest improvements:

GitHub Issue (Preferred) 📧 Email (No GitHub)
← Back to All Events