GEM Training: How Meta Doubled the Efficiency of Its LLM-Scale Ads Foundation Model
中文摘要
Meta通过优化GEM模型,将LLM规模下的广告推荐训练效率翻倍,实现了高效的大规模扩展。
English Summary
Meta doubled the training efficiency of its GEM ads recommendation model, enabling efficient LLM-scale training for Instagram and Facebook.
Original Excerpt
Meta’s Generative Ads Recommendation Model (GEM), the foundation model behind ads recommendations across Instagram and Facebook, now trains at LLM scale on several thousand of the latest-generation GPUs. This post goes into the details on how we achieved: doubling end-to-end (E2E) training efficiency to 20–25% Model FLOPs Utilization (MFU) while scaling training FLOPs 4x in [...] Read More... The post GEM Training: How Meta Doubled the Efficiency of Its LLM-Scale Ads Foundation Model appeared first on Engineering at Meta.