Back to Home
AI on Medium··Industry Media

On-Policy vs Off-Policy Learning: The Most Misunderstood Distinction in Reinforcement Learning

中文摘要

本文详细解释了强化学习中 On-Policy 与 Off-Policy 的区别,涵盖从 TD 误差到 GRPO 的概念,并在文末提供代数推导。

English Summary

This article explains the distinction between on-policy and off-policy reinforcement learning, covering concepts from TD errors to GRPO with algebraic derivations at the end.

Original Excerpt

From TD errors to GRPO — explained in words, with all the algebra kept in one place at the end, Continue reading on Medium »