📜

Writeup: Progress on AI Safety via Debate

📅 2020
policy development
⚪ Common
#ai #debate (ai safety technique) #factored cognition #iterated amplification

📖 Description

This is a writeup of the research done by the "Reflection-Humans" team at OpenAI in Q3 and Q4 of 2019. During that period we investigated mechanisms that would allow evaluators to get correct and helpful answers from experts, without the evaluators themselves being expert in the domain of the questions. This follows from the original work on [AI Safety via Debate](https://arxiv.org/abs/1805.00899) and the [call for research on human aspects of AI safety](https://distill.pub/2019/safety-needs-social-scientists/ ), and is also closely related to work on [Iterated Amplification](https://openai.com/blog/amplifying-ai-training/).

📊 Game Impacts

Variable Change Condition
Research +10 Always
Vibey Doom +2 Always
Ethics Risk -5 Always

💭 Reactions

🔬 Safety Researcher Reaction: ⚠️ Placeholder - Needs Real Quote
"Adds to our knowledge base"
📰 Media Reaction: ⚠️ Placeholder - Needs Real Quote
"Discussed in AI safety community"
💡 Found a Real Quote? Suggest it here

🔗 Sources

🏷️ Event Metadata

Think this event's metadata could be improved? Suggest changes to category, rarity, tags, game impacts, or p(doom) effects.

🤝 Found an Issue?

This event data is sourced from the pdoom-data repository. If you notice errors or want to suggest improvements:

GitHub Issue (Preferred) 📧 Email (No GitHub)
← Back to All Events