Back to Home
AI on Medium··Industry Media

On-Policy Distillation: How Smaller LLMs Learn From Their Own Mistakes

中文摘要

在策略蒸馏通过让小型大语言模型从自身错误中学习,使其在追求规模的同时,也能显著提升运行效率。

English Summary

On-policy distillation helps smaller LLMs improve efficiency by allowing them to learn from their own mistakes.

Original Excerpt

AI models are not only becoming larger. They are also becoming more efficient. Continue reading on Medium »