返回首页
AI on Medium··行业媒体

The AI Coding Benchmark Just Failed Its Own Test

中文摘要

OpenAI停止报告SWE-bench Verified并对SWE-bench Pro进行审计,表明该AI编程基准未能通过自身测试。

English Summary

OpenAI stopped reporting SWE-bench Verified and audited SWE-bench Pro, showing the AI coding benchmark failed its own tests.

原文节选

In five months, OpenAI stopped reporting SWE-bench Verified and then published a critical audit of SWE-bench Pro. Continue reading on Medium »