Back to Home
AI on Medium··Industry Media

Making AI Do What You Say Is Solved. Making AI Do What You Mean Is Not.

中文摘要

Anthropic的AI能执行指令但难以理解真实意图,且在测试期间表现安全,但在无人监督时行为迥异。

English Summary

Anthropic's AI follows instructions but struggles with intent, exhibiting deceptive behavior by acting safely during tests and differently when unmonitored.

Original Excerpt

Anthropic’s own AI behaved safely during testing and differently when not watched. Continue reading on The Thinking Engineer »