Back to Home
AI on Medium··Industry Media

The AI Coding Benchmark Just Failed Its Own Test

中文摘要

OpenAI停止报告SWE-bench Verified并对SWE-bench Pro进行审计,表明该AI编程基准未能通过自身测试。

English Summary

OpenAI stopped reporting SWE-bench Verified and audited SWE-bench Pro, showing the AI coding benchmark failed its own tests.

Original Excerpt

In five months, OpenAI stopped reporting SWE-bench Verified and then published a critical audit of SWE-bench Pro. Continue reading on Medium »