返回首页
AI on Medium··行业媒体

Unexpressible, Not Filtered: A Structural Approach to Agent Safety

中文摘要

作者建议通过结构化设计使AI智能体无法表达不安全行为,而非单纯依靠过滤,以实现更本质的安全保障。

English Summary

Instead of filtering unsafe AI actions, the author proposes a structural approach that makes certain unsafe behaviors impossible to express, ensuring fundamental agent safety.

原文节选

Why I stopped trying to catch unsafe AI-agent actions and made a class of them impossible to express instead — and the honest limits of… Continue reading on Medium »