返回首页
RadarAI··论文与技术

【Agentic RL / 强化学习 / OPD】Hermes & OPD 源码阅读笔记

中文摘要

本文对比OpenClaw-RL与Hermes源码,解析了基于后见之明提示的在线策略蒸馏(OPD)机制及其工程实现。

English Summary

This article compares OpenClaw-RL and Hermes source code to analyze the implementation of On-Policy Distillation (OPD) using hindsight prompting and LLM feedback.

原文节选

📌 一句话摘要 本文通过对 OpenClaw-RL 和 Hermes 源码的对比阅读,深度解析了基于后见之明提示的在线策略蒸馏(OPD)机制及其工程实现。 📝 详细摘要 文章详细拆解了 OPD(On-Policy Distillation)的核心工作流:通过 LLM Judge 从环境反馈中提取后验知识(Hint),构建增强 Prompt,利用 vLLM 的 `prompt_logprobs` 实现 Teacher S...