Passed the Jailbreak Test. That Wasn’t the Deployment Test.
中文摘要
模型级越狱测试并不等同于部署安全。对抗性评估无法确保 AI 代理在使用工具时的实际运行安全。
English Summary
Model-level jailbreak tests do not guarantee deployment safety. Resistance to adversarial prompts cannot ensure an AI agent is safe when using tools.
Original Excerpt
Model-level evaluations can measure resistance to adversarial prompts. They cannot decide whether an agent is safe to operate with tools… Continue reading on Medium »