Benchmark for successful concept extrapolation/avoiding goal misgeneralization
📖 Description
If an AI has been trained on data about adults, what should it do when it encounters a child? When an AI encounters a situation that its training data hasn't covered, there's a risk that it will incorrectly generalize from what it has been trained for and do the wrong thing. [Aligned AI](https://buildaligned.ai/news/aligned-ai-releases-new-disambiguation-benchmark/) has released a new [benchmark](https://github.com/alignedai/HappyFaces) designed to measure how well image-classifying algorithms avoid goal misgeneralization. This post explains what goal misgeneralization is, what the new benchmark measures, and how the benchmark is related to goal misgeneralization and concept extrapolation.
📊 Game Impacts
| Variable | Change | Condition |
|---|---|---|
| Research | +5 | Always |
💭 Reactions
"Background research"
"Discussed in AI safety community"
🤝 Found an Issue?
This event data is sourced from the pdoom-data repository. If you notice errors or want to suggest improvements:
GitHub Issue (Preferred) 📧 Email (No GitHub)