从 KL 的方向看 SFT 与 RL:大模型到底是在”学会做”,还是在”学会选”?
中文摘要
文章分析KL散度方向性,指出SFT对应覆盖分布的Forward KL,而RL则通过Reverse KL聚焦高奖励,探讨了模型在后训练中的学习机制。
English Summary
The article analyzes KL divergence directions, explaining how SFT covers distributions via Forward KL while RL focuses on high-reward areas through Reverse KL during post-training.
Original Excerpt
📌 一句话摘要 本文从 KL 散度的方向性出发,深入分析了 SFT 对应 Forward KL(覆盖目标分布)、RL/RLHF 对应 Reverse KL(聚焦高奖励区域)的数学原理与训练行为差异。 📝 详细摘要 本文从 KL 散度的方向性出发,系统阐述了 Forward KL 与 Reverse KL 在大模型后训练中的对应关系。文章首先解释了 KL 散度非对称性的根源在于期望在哪个分布上取,进而指出 SFT 本质上...