Back to Home
量子位··Domestic Sources

机器人不能停下来等模型:星尘发布 SmoothRL,让在线强化学习跟上大模型的异步推理

中文摘要

星尘智能发布SmoothRL框架,通过异步执行让在线强化学习速度匹配大模型的推理效率。

English Summary

Astribot released SmoothRL, an asynchronous online reinforcement learning framework designed to match large model inference speeds.

Original Excerpt

星尘智能(Astribot) 基座模型团队发布能异步执行的在线强化学习框架 SmoothRL