GRAM: Anthropic’s New Way to Make AI Forget Dangerous Knowledge Without Retraining the Entire Model
中文摘要
Anthropic 正在探索 GRAM 技术,通过控制模型的学习内容而非重新训练整个模型,使 AI 能够“忘记”危险知识。
English Summary
Anthropic is exploring GRAM, a method to make AI forget dangerous knowledge by controlling what it learns instead of retraining the entire model.
原文节选
Instead of teaching AI what not to say, Anthropic is exploring something much deeper: controlling what the model actually learns. Continue reading on Medium »