AI Safety Needs Social Scientists
📖 Description
The goal of long-term artificial intelligence (AI) safety is to ensure that advanced AI systems are reliably aligned with human values???that they reliably do things that people want them to do.Roughly by human values we mean whatever it is that causes people to choose one option over another in each case, suitably corrected by reflection, with differences between groups of people taken into account. There are a lot of subtleties in this notion, some of which we will discuss in later sections and others of which are beyond the scope of this paper. Since it is difficult to write down precise rules describing human values, one approach is to treat aligning with human values as another learning problem. We ask humans a large number of questions about what they want, train an ML model of their values, and optimize the AI system to do well according to the learned values.
📊 Game Impacts
| Variable | Change | Condition |
|---|---|---|
| Research | +15 | Always |
| Papers | +10 | Always |
| Vibey Doom | +5 | Always |
💭 Reactions
"Notable work on AI safety"
"Visual machine learning research"
🤝 Found an Issue?
This event data is sourced from the pdoom-data repository. If you notice errors or want to suggest improvements:
GitHub Issue (Preferred) 📧 Email (No GitHub)