Back to Home
RadarAI··Papers & Tech

Claude 通过率不到 4%,SaaS-Bench 撕碎了 Computer-Use 的「全自动办公」幻想

中文摘要

SaaS-Bench测试显示Claude在SaaS长程任务中通过率仅3.8%,揭示了当前Computer-Use Agent难以实现全自动办公。

English Summary

SaaS-Bench reveals Claude's 3.8% pass rate in SaaS tasks, debunking the illusion of fully automated office workflows via Computer-Use Agents.

Original Excerpt

📌 一句话摘要 SaaS-Bench 基准测试通过 23 个真实 SaaS 系统、106 个长程任务,揭示了当前最强 Computer-Use Agent 在真实办公场景中的惨淡表现:Claude Opus 4.7 端到端通过率仅 3.8%,暴露了 Agent 在长程任务中的四种结构性失败模式。 📝 详细摘要 本文介绍了 UniPat AI 团队推出的 SaaS-Bench 基准测试,旨在评估 AI Agent 在真实...