返回首页
AI on Medium··行业媒体

Evals: I Let an AI Grade Another AI. Here’s Why That’s Risky

中文摘要

作者探讨了使用AI评估另一个AI的潜在风险,以及如何对输出不一致的大语言模型进行测试。

English Summary

The author discusses the risks of using AI to evaluate other AI models and how to test non-deterministic LLMs.

原文节选

In my last post I showed how you test an LLM that never gives the same answer twice. The trick was to run each case many times and treat… Continue reading on Medium »