The AI Coding Benchmark Just Failed Its Own Test
中文摘要
OpenAI停止报告SWE-bench Verified并对SWE-bench Pro进行审计,表明该AI编程基准未能通过自身测试。
English Summary
OpenAI stopped reporting SWE-bench Verified and audited SWE-bench Pro, showing the AI coding benchmark failed its own tests.
Original Excerpt
In five months, OpenAI stopped reporting SWE-bench Verified and then published a critical audit of SWE-bench Pro. Continue reading on Medium »