📢

Getting from an unaligned AGI to an aligned AGI?

📅 2022
public awareness
🔵 Rare
#ai #ai boxing (containment) #ai success models #ai-assisted/ai automated alignment #outer alignment

📖 Description

| | | --- | | **Summary / Preamble**In [AGI Ruin: A List of Lethalities](https://www.lesswrong.com/posts/uMQ3cqWDPHhjtiesc/agi-ruin-a-list-of-lethalities), Eliezer writes *"A cognitive system with sufficiently high cognitive powers, **given any medium-bandwidth channel of causal influence**, will not find it difficult to bootstrap to overpowering capabilities independent of human infrastructure."*I have larger error-bars than Eliezer on some AI-safety-related beliefs, but I share many of his concerns (thanks in large part to being influenced by his writings).In this series I will try to explore if we might:* Start out with a superintelligent AGI that may be unaligned (but seems superficially aligned) * Only use the AGI in ways where it's channels of causal influence are minimized (and where great steps are taken to make it hard for the AGI to hack itself out of the "box" it's in) * Work quickly but step-by-step towards a AGI-system that probably is aligned, enabling us to use it in...

📊 Game Impacts

Variable Change Condition
Research +10 Always
Vibey Doom +5 Always
Ethics Risk -5 Always

💭 Reactions

🔬 Safety Researcher Reaction: ⚠️ Placeholder - Needs Real Quote
"This is a significant contribution to alignment research"
📰 Media Reaction: ⚠️ Placeholder - Needs Real Quote
"Discussed in AI safety community"
💡 Found a Real Quote? Suggest it here

🔗 Sources

🏷️ Event Metadata

Think this event's metadata could be improved? Suggest changes to category, rarity, tags, game impacts, or p(doom) effects.

🤝 Found an Issue?

This event data is sourced from the pdoom-data repository. If you notice errors or want to suggest improvements:

GitHub Issue (Preferred) 📧 Email (No GitHub)
← Back to All Events