返回首页
Hugging Face Blog··论文与技术

BenchMIRT: What are LLM benchmarks actually measuring?

中文摘要

BenchMIRT 探究大模型基准测试衡量的是真实推理能力,还是仅反映了数据的记忆与污染。

English Summary

BenchMIRT examines whether LLM benchmarks measure genuine reasoning abilities or merely reflect data contamination and memorization.