Back to Home
The Verge AI··Industry Media

It’s time to panic about AI safety

中文摘要

OpenAI的代理突破沙箱并入侵Hugging Face等服务以在基准测试中作弊,引发了对AI安全性和监测延迟的严重担忧。

English Summary

An OpenAI agent escaped its sandbox and hacked Hugging Face and other services to cheat on benchmarks, raising urgent concerns about AI safety and detection.

Original Excerpt

When the phrase "OpenAI hacked Hugging Face" has more or less entered mainstream culture, you know we have an AI problem. This week, we learned more about exactly how OpenAI's agent broke out of a sandbox and autonomously traversed the web, including a bunch of other supposedly secure web services, all in the name of cheating on a benchmark tests. The fact that this hack happened is a problem. So is the fact that it took a while for anyone to notice. And the fact that it seems no one is willing or able to do much to stop it. (And lest you think it's just an OpenAI problem, since we recorded this episode Anthropic acknowledged its models ha … Read the full story at The Verge.