Evals: I Let an AI Grade Another AI. Here’s Why That’s Risky
中文摘要
作者探讨了使用AI评估另一个AI的潜在风险,以及如何对输出不一致的大语言模型进行测试。
English Summary
The author discusses the risks of using AI to evaluate other AI models and how to test non-deterministic LLMs.
Original Excerpt
In my last post I showed how you test an LLM that never gives the same answer twice. The trick was to run each case many times and treat… Continue reading on Medium »