返回首页
RadarAI··论文与技术

“智能体最后的考试”,Fable 5 竟然不敌 GPT 5.5

中文摘要

UC Berkeley发布ALE测试,GPT 5.5微弱胜过Fable 5,但在极难任务中通过率均不足3%。

English Summary

UC Berkeley's ALE benchmark shows GPT 5.5 narrowly beating Claude Fable 5, with all models achieving under 3% success on difficult tasks.

原文节选

📌 一句话摘要 UC Berkeley 发布全新 Agent 基准测试 ALE,在真实专业软件操作任务中,GPT 5.5 以微弱优势击败 Claude Fable 5,所有模型在最难任务上几乎全军覆没,通过率不足 3%。 📝 详细摘要 UC Berkeley 团队发布名为「Agents' Last Exam」(ALE)的全新 AI Agent 基准测试,旨在评估 AI 在真实专业软件环境中的实操能力。测试涵盖 Siem...