Building a Sparse Autoencoder on GPT-2 from Scratch: A Mechanistic Interpretability Investigation
中文摘要
本研究通过在 GPT-2 上从零构建稀疏自编码器,探究机械解释性以揭示模型的内部机制。
English Summary
This research explores mechanistic interpretability by building a sparse autoencoder on GPT-2 to investigate the internal workings of language models.
Original Excerpt
What happens when you wiretap a language model’s brain Continue reading on Medium »