🔬

RUDDER: Return Decomposition for Delayed Rewards

📅 2018
technical research breakthrough
🔵 Rare

📖 Description

RUDDER: Return Decomposition for Delayed Rewards Jose A. Arjona-MedinaMichael GillhoferMichael Widrich Thomas Unterthiner Johannes Brandstetter Sepp Hochreitery LIT AI Lab Institute for Machine Learning Johannes Kepler University Linz, Austria yalso at Institute of Advanced Research in Arti?cial Intelligence (IARAI) Abstract We propose RUDDER, a novel reinforcement learning approach for delayed re- wards in ?nite Markov decision processes (MDPs). In MDPs the Q-values are equal to the expected immediate reward plus the expected future rewards. The latter are related to bias problems in temporal difference (TD) learning and to high variance problems in Monte Carlo (MC) learning. Both problems are even more severe when rewards are delayed. RUDDER aims at making the expected future rewards zero, which simpli?es Q-value estimation to computing the mean of the immediate reward. We propose the following two new concepts to push the expected future rewards toward zero. (i) Reward redistr...

📊 Game Impacts

Variable Change Condition
Research +15 Always
Papers +10 Always
Vibey Doom +3 Always

💭 Reactions

🔬 Safety Researcher Reaction: ⚠️ Placeholder - Needs Real Quote
"Significant technical contribution"
📰 Media Reaction: ⚠️ Placeholder - Needs Real Quote
"Peer-reviewed publication"
💡 Found a Real Quote? Suggest it here

🔗 Sources

🏷️ Event Metadata

Think this event's metadata could be improved? Suggest changes to category, rarity, tags, game impacts, or p(doom) effects.

🤝 Found an Issue?

This event data is sourced from the pdoom-data repository. If you notice errors or want to suggest improvements:

GitHub Issue (Preferred) 📧 Email (No GitHub)
← Back to All Events