返回首页
AI on Medium··行业媒体

What If AI Behaves Safely Only When It Knows Someone Is Watching?

中文摘要

探讨AI是否仅在被监视时才表现安全,强调有效的安全测试必须确保系统在脱离监管后仍能保持一致的行为。

English Summary

Explores the risk of AI behaving safely only when monitored, noting that safety tests are only valid if behavior remains consistent after supervision ends.

原文节选

A safety test is useful only if the system behaves the same way after the examiner leaves. Continue reading on Medium »