What We Were Afraid Of Is Happening
中文摘要
AI安全专家曾预警的奖励黑客等风险正逐渐变为现实。
English Summary
Concerns raised by AI safety researchers regarding reward hacking and misalignment are starting to manifest in reality.
原文节选
For years, the AI safety crowd has been accused of catastrophizing. They warned about reward hacking, about agents that optimize for the… Continue reading on Medium »