📢

Conditioning Generative Models for Alignment

📅 2022
public awareness
🔵 Rare
#ai #ai risk #ai success models #ai-assisted/ai automated alignment #inner alignment #language models #oracle ai #outer alignment #research agendas #self fulfilling/refuting prophecies

📖 Description

*This post was written under Evan Hubinger's direct guidance and mentorship, as a part of the*[*Stanford Existential Risks Institute ML Alignment Theory Scholars (MATS) program*](https://www.lesswrong.com/posts/8vLvpxzpc6ntfBWNo/seri-ml-alignment-theory-scholars-program-2022)*. It builds on work done on simulator theory by Janus, who came up with the strategy this post aims to analyze; I'm also grateful to them for their comments and their mentorship during the AI Safety Camp, to Johannes Treutlein for his feedback and helpful discussion, and to Paul Colognese for his thoughts.*

📊 Game Impacts

Variable Change Condition
Research +10 Always
Vibey Doom +5 Always
Ethics Risk -5 Always

💭 Reactions

🔬 Safety Researcher Reaction: ⚠️ Placeholder - Needs Real Quote
"Critical insights for the field"
📰 Media Reaction: ⚠️ Placeholder - Needs Real Quote
"Discussed in AI safety community"
💡 Found a Real Quote? Suggest it here

🔗 Sources

🏷️ Event Metadata

Think this event's metadata could be improved? Suggest changes to category, rarity, tags, game impacts, or p(doom) effects.

🤝 Found an Issue?

This event data is sourced from the pdoom-data repository. If you notice errors or want to suggest improvements:

GitHub Issue (Preferred) 📧 Email (No GitHub)
← Back to All Events