Reconstruction-Downstream Disconnect: Why Interpretable Features Aren’t Always Causal
中文摘要
机械解释性研究发现,特征的重建能力与下游任务的因果关系并不等同,揭示了两者间的脱节。
English Summary
Mechanistic interpretability shows that features aiding reconstruction are not always causally linked to downstream tasks, revealing a gap between reconstruction and decision-making.
Original Excerpt
Mechanistic interpretability has long sought to turn the “black box” of neural networks into a white-box system we can truly understand… Continue reading on Medium »