Answering questions honestly instead of predicting human answers: lots of problems and some solutions
📖 Description
*This post is the result of work I did with Paul Christiano on the ideas in his "[Teaching ML to answer questions honestly instead of predicting human answers](https://www.alignmentforum.org/posts/QqwZ7cwEA2cxFEAun/teaching-ml-to-answer-questions-honestly-instead-of)" post. In addition to expanding upon what is in that post in terms of identifying numerous problems with the proposal there and identifying ways in which some of those problems can be patched, I think that this post also provides a useful window into what Paul-style research looks like from a non-Paul perspective.*
📊 Game Impacts
| Variable | Change | Condition |
|---|---|---|
| Research | +5 | Always |
| Vibey Doom | +2 | Always |
| Ethics Risk | -5 | Always |
💭 Reactions
🔬 Safety Researcher Reaction:
⚠️ Placeholder - Needs Real Quote
"Useful research for the community"
"Useful research for the community"
📰 Media Reaction:
⚠️ Placeholder - Needs Real Quote
"Discussed in AI safety community"
💡 Found a Real Quote? Suggest it here
"Discussed in AI safety community"
🤝 Found an Issue?
This event data is sourced from the pdoom-data repository. If you notice errors or want to suggest improvements:
GitHub Issue (Preferred) 📧 Email (No GitHub)