Aller au contenu
OpenAI News·· 2018-10-31sélectionAI Score62

OpenAI 提出基于预测奖励的强化学习方法 RND

Reinforcement learning with prediction-based rewards

AI Introduction

OpenAI 开发了 Random Network Distillation (RND), 一种基于预测的奖励方法, 通过好奇心鼓励强化学习智能体探索环境, 在 Montezuma's Revenge 上首次超过人类平均表现.

Source :OpenAI News · openai.com