EMO: Pretraining mixture of experts for emergent modularity
中文摘要
EMO通过预训练混合专家模型实现突现模块化,在不增加计算成本的情况下,显著提升了模型的专业化能力与任务表现。
English Summary
EMO shows that pretraining Mixture-of-Experts leads to emergent modularity, significantly enhancing model specialization and performance without increasing computational overhead during inference.