Is it agentic enough? Benchmarking open models on your own tooling
中文摘要
本文通过在自定义工具上进行基准测试,评估开源模型作为智能体的实际表现。
English Summary
This study benchmarks open-source models to evaluate their performance as agents when integrated with custom user-provided tools.