Decoding Transformers, One Equation at a Time: Multi-Head Attention and Where Attention Is Used
中文摘要
Transformer架构系列第二部分,聚焦多头注意力机制及其应用,以公式化方式简化理解。
English Summary
Part 2 of a series decoding Transformer architecture. It focuses on Multi-Head Attention and where attention is used, simplifying concepts with equations.
Original Excerpt
Part 2 of a learning-in-public series on the Transformer architecture Continue reading on Medium »