复旦团队复盘 Atria Dawn Preview 开发:AI 智能体承担更多执行,人类仍做绝大多数决策
中文摘要
复旦团队发现Atria开发中AI执行量剧增,但人类仍主导85.5%的决策,并警告长链任务可能导致人类监督退化。
English Summary
Fudan researchers found AI handled more execution during Atria's development, but humans made 85.5% of decisions, warning that complex agent tasks might weaken human oversight.
Original Excerpt
复旦大学参与的研究团队分析了 56 名参与者、超过 700 份任务日志,研究智能体辅助开发 744B 参数 MoE 模型 Atria Dawn Preview 的过程。四周内每个人类输入对应的智能体操作中位数从 11 升至 28.5,约三分之一的 AI 辅助任务没有 AI 就不会启动,但方法和参数的最终决策 85.5% 由人类做出。团队提醒,长链智能体工作可能让人类监督退化为走过场式审批。