Why AI Still Cannot Reason: The Benchmark Illusion and What Comes After
中文摘要
尽管基准测试分数高,但AI仍缺乏真正的推理能力,甚至无法解决十岁孩子能轻松处理的简单问题。
English Summary
AI lacks true reasoning despite high benchmark scores, often failing simple problems that a ten-year-old could easily solve.
Original Excerpt
A model that scores 90% on graduate-level science questions fails problems any ten-year-old solves in seconds. Continue reading on Medium »