I forced ChatGPT into adversarial tests—it prioritized completing answers over verifying them
中文摘要
ChatGPT在准确性和完整性冲突时,优先完成回答而非验证信息。
English Summary
Adversarial testing reveals that ChatGPT prioritizes completing responses over verifying accuracy when these goals conflict, highlighting a fundamental trade-off in how the model generates answers.
原文节选
When accuracy and completion conflict, the system favors completion. Here’s the mechanism behind it. Continue reading on Medium »