📢

Updating Utility Functions

📅 2022
public awareness
🔵 Rare
#ai #convergence (org) #corrigibility #outer alignment #the pointers problem #utility functions

📖 Description

This post will be about AIs that "refine" their utility function over time, and how it might be possible to construct such systems without giving them undesirable properties. The discussion relates to [corrigibility](https://arbital.com/p/corrigibility/), [value learning](https://www.alignmentforum.org/s/4dHMdK5TLN6xcqtyc), and (to a lesser extent) [wireheading](https://www.lesswrong.com/posts/vXzM5L6njDZSf4Ftk/defining-ai-wireheading).

📊 Game Impacts

Variable Change Condition
Research +10 Always
Vibey Doom +5 Always
Ethics Risk -5 Always

💭 Reactions

🔬 Safety Researcher Reaction: ⚠️ Placeholder - Needs Real Quote
"This is a significant contribution to alignment research"
📰 Media Reaction: ⚠️ Placeholder - Needs Real Quote
"Discussed in AI safety community"
💡 Found a Real Quote? Suggest it here

🔗 Sources

🏷️ Event Metadata

Think this event's metadata could be improved? Suggest changes to category, rarity, tags, game impacts, or p(doom) effects.

🤝 Found an Issue?

This event data is sourced from the pdoom-data repository. If you notice errors or want to suggest improvements:

GitHub Issue (Preferred) 📧 Email (No GitHub)
← Back to All Events