Back to Home
RadarAI··Papers & Tech

ICML 2026 | AI Scientist 真会科学发现吗?30 个前沿模型无一过线

中文摘要

ICML 2026论文CausalGame揭示,30款前沿LLM Agent在因果推理上普遍失效,表明AI Scientist仍缺乏真正的科学发现能力。

English Summary

ICML 2026 paper CausalGame shows 30 leading LLM agents fail at causal reasoning, indicating that "AI Scientists" still lack true scientific discovery capabilities.

Original Excerpt

📌 一句话摘要 ICML 2026 Oral 论文 CausalGame 揭示,30 款前沿 LLM Agent 在因果推理上普遍失效,凸显 AI Scientist 仍缺乏真正科学发现能力。 📝 详细摘要 文章介绍了由 MBZUAI、卡内基梅隆大学等机构提出的 CausalGame 基准,该基准通过结构因果模型(SCM)驱动的交互式游戏评估 LLM Agent 在因果推理、实验设计和机制解释方面的能力。评测覆盖 Op...