BenchMIRT: What are LLM benchmarks actually measuring?
中文摘要
BenchMIRT 探究大模型基准测试衡量的是真实推理能力,还是仅反映了数据的记忆与污染。
English Summary
BenchMIRT examines whether LLM benchmarks measure genuine reasoning abilities or merely reflect data contamination and memorization.