Eight models, five vendors, one answer sheet
中文摘要
测试显示五家厂商的八款旗舰AI编码模型仅能解决3/10的新任务,且在其余部分存在误导性回答。
English Summary
Tests on eight flagship coding models from five vendors revealed they only solved 3/10 new tasks and hallucinated the rest.
Original Excerpt
I spent $66 grading AI coding models with tests they couldn't see. Every flagship solved the same 3/10 fresh tasks and lied about the rest. Continue reading on Medium »