An Anthropic researcher just gave us a peek at self-improving AI
中文摘要
Anthropic研究员展示了自我提升的AI,在不降低整体性能的情况下,成功修复了10项对齐基准测试中的行为问题。
English Summary
An Anthropic researcher demonstrated self-improving AI that corrected 10 misaligned behavior benchmarks without compromising overall performance.
原文节选
Given 10 benchmarks for specific misaligned behaviors, the automated systems were able to improve performance on every single one without degrading overall performance.