Pre-training Under Infinite Compute: Rethinking Data Efficiency When Tokens Become Scarce
中文摘要
本文探讨在计算资源充足但高质量Token短缺的情况下,如何重新思考大模型预训练的数据效率。
English Summary
This article explores rethinking data efficiency for LLM pre-training in scenarios where compute is abundant but high-quality tokens are scarce.
Original Excerpt
1. Why I Care About This Problem Continue reading on Medium »