返回首页
AI on Medium··行业媒体

Coding Benchmarks Are Dying. MirrorCode Asks AI to Rebuild the Whole Program.

中文摘要

传统代码基准测试已过时,MirrorCode 提出让 AI 重构整个程序的新方法,以更好地评估前沿智能体的能力。

English Summary

Traditional coding benchmarks are obsolete; MirrorCode proposes having AI rebuild entire programs to better evaluate the capabilities of frontier agents.

原文节选

Programming tests used to measure whether AI could solve a difficult coding problem. Now frontier agents can work for days, spend billions… Continue reading on Medium »