DropoutDAgger: A Bayesian Approach to Safe Imitation Learning
📖 Description
arXiv:1709.06166v1 [cs.AI] 18 Sep 2017DropoutDAgger: A Bayesian Approach to Safe Imitation Learn ing Kunal Menda, Katherine Driggs-Campbell, and Mykel J. Koche nderfer Abstract ? While imitation learning is becoming com- mon practice in robotics, this approach often suffers from data mismatch and compounding errors. DAgger is an iterative algorithm that addresses these issues by continually aggregating training data from both the expert and novice policies, but does not consider the impact of safety. We present a probabilistic extension to DAgger, which uses the distribution over actions provided by the novice policy, for a given observation. Our method, which we call DropoutDAgger, uses dropout to train the novice as a Bayesian neural network that provides insight to its con?dence. Using the distribution over the novice?s action s, we estimate a probabilistic measure of safety with respect to the expert action, tuned to balance exploration and exploitation. The utility of this ap...
📊 Game Impacts
| Variable | Change | Condition |
|---|---|---|
| Research | +15 | Always |
| Papers | +10 | Always |
| Vibey Doom | +3 | Always |
💭 Reactions
"Advances our understanding of AI safety"
"Published in academic venue"
🔗 Sources
🤝 Found an Issue?
This event data is sourced from the pdoom-data repository. If you notice errors or want to suggest improvements:
GitHub Issue (Preferred) 📧 Email (No GitHub)