Behavior is not Evidence
中文摘要
探讨AI控制的局限性,指出仅凭观察AI的行为输出不足以证明其内部安全性或可靠性。
English Summary
Discusses the limitations of AI control, arguing that observing external behavior is insufficient evidence of a system's actual safety or internal mechanisms.
Original Excerpt
Every meaningful control we have built for AI depends on watching what the AI does. An AI system is put through an evaluation, its outputs… Continue reading on Predict »