Back to Home
美团技术团队··Domestic Sources

美团 LongCat 开源 General 365:树立推理评测新标尺

中文摘要

美团 LongCat 开源 General 365 推理评测基准。实测显示即便 Gemini 3 Pro 准确率也仅 62.8%,多数模型未能及格。

English Summary

Meituan LongCat released General 365 reasoning benchmark. Tests show even Gemini 3 Pro only achieves 62.8% accuracy, while most models fail to reach 60%.

Original Excerpt

美团 LongCat 团队正式发布 General 365。我们发现,在对 26 款主流模型的实测中,目前地表最强的 Gemini 3 Pro 准确率仅为 62.8%,而绝大多数模型甚至没能摸到 60 分的及格线。