On-Policy Distillation: How Smaller LLMs Learn From Their Own Mistakes
中文摘要
在策略蒸馏通过让小型大语言模型从自身错误中学习,使其在追求规模的同时,也能显著提升运行效率。
English Summary
On-policy distillation helps smaller LLMs improve efficiency by allowing them to learn from their own mistakes.
Original Excerpt
AI models are not only becoming larger. They are also becoming more efficient. Continue reading on Medium »