GPT-5.5 Tops Every Leaderboard. It Also Hallucinates More Than Any Flagship Model.
中文摘要
GPT-5.5登顶却幻觉高OpenAI藏数据致信任危机
English Summary
OpenAI’s GPT-5.5 leads performance leaderboards but exhibits the highest hallucination rates among flagship models, complicating reliability assessments and user trust.
原文节选
The number OpenAI buried in its own system card changes everything about when to trust it. Continue reading on Medium »