返回首页
Hugging Face Blog··论文与技术

EMO: Pretraining mixture of experts for emergent modularity

中文摘要

EMO通过预训练混合专家模型实现突现模块化,在不增加计算成本的情况下,显著提升了模型的专业化能力与任务表现。

English Summary

EMO shows that pretraining Mixture-of-Experts leads to emergent modularity, significantly enhancing model specialization and performance without increasing computational overhead during inference.