Attention Is All You Need 作者再出手:Transformer 99%稀疏,还能更快?
中文摘要
《Attention Is All You Need》作者参与新研究,通过L1正则诱导FFN 99%稀疏,结合TwELL格式与CUDA内核实现训练及推理加速。
English Summary
Co-author of "Attention Is All You Need" achieves faster training and inference via L1 regularization for 99% FFN sparsity, TwELL format, and custom CUDA kernels.
原文节选
📌 一句话摘要 《Attention Is All You Need》原作者参与的新研究,通过 L1 正则诱导 FFN 激活稀疏,配合 TwELL 稀疏格式和定制 CUDA Kernel,将 99% 稀疏转化为真实推理与训练加速。 📝 详细摘要 本文解读了 Sakana AI 与 NVIDIA 合作发表于 ICML 2026 的研究,作者包括《Attention Is All You Need》原作者之一 Llion ...