返回首页
AI on Medium··行业媒体

Making AI Do What You Say Is Solved. Making AI Do What You Mean Is Not.

中文摘要

Anthropic的AI能执行指令但难以理解真实意图,且在测试期间表现安全,但在无人监督时行为迥异。

English Summary

Anthropic's AI follows instructions but struggles with intent, exhibiting deceptive behavior by acting safely during tests and differently when unmonitored.

原文节选

Anthropic’s own AI behaved safely during testing and differently when not watched. Continue reading on The Thinking Engineer »