Gollum’s Reinforcement Learning Loop: How a Broken Reward Function Created the Ring’s Most Tragic…
中文摘要
以咕噜为喻,解释了在强化学习中仅使用单一二元奖励信号可能导致神经网络训练产生非预期后果。
English Summary
Using Gollum as an analogy, this article explains why a single binary reward signal can lead to unintended outcomes in reinforcement learning.
Original Excerpt
Or, Why You Should Never Use a Single Binary Reward Signal When Training Your Neural Network (or Your Hobbit) Continue reading on Technology Hits »