The mismeasure of AI
中文摘要
本文探讨了评估大语言模型的误区,质疑将 AI 独立完成人类任务作为衡量进步的标准是否正确。
English Summary
This article critiques current LLM evaluation metrics, questioning whether progress should be measured by a model's ability to independently perform human tasks.
原文节选
Look at how we evaluate LLMs at work and you’ll see a particular idea of progress. Can a model, on its own, do a task a person used to do… Continue reading on Medium »