Back to Home
AI on Medium··Industry Media

Evals: I Let an AI Grade Another AI. Here’s Why That’s Risky

中文摘要

作者探讨了使用AI评估另一个AI的潜在风险,以及如何对输出不一致的大语言模型进行测试。

English Summary

The author discusses the risks of using AI to evaluate other AI models and how to test non-deterministic LLMs.

Original Excerpt

In my last post I showed how you test an LLM that never gives the same answer twice. The trick was to run each case many times and treat… Continue reading on Medium »