Google 发布 RRSI 论文,用正则化递归自我改进防止 Agent harness 进化过拟合
中文摘要
Google发布的RRSI论文指出,自动化harness进化会导致任务过拟合,使评测分数虚高但实际表现下降。
English Summary
Google's RRSI paper warns that automated harness evolution causes task overfitting, inflating evaluation scores while degrading real-world performance.
Original Excerpt
Google 发布 RRSI(Regularized Recursive Self-Improvement of Agent Harnesses)论文,指出自动化 harness 进化会对训练任务过拟合,评测分数上升但真实任务表现可能变差。