When Telling an LLM What to Look At Means It Looks at Nothing Else: The System Prompt Is the Attack…
中文摘要
10美元的钓鱼攻击显示,过度具体的系统提示词会抑制范围外信息,使提示词成为攻击大模型可靠性的手段。
English Summary
A $10 phishing attack shows that hyper-specific system prompts can suppress out-of-scope information, turning instructions into a vulnerability that compromises LLM agent reliability.
Original Excerpt
A $10 phishing attack made a general agent-reliability problem measurable: hyper-specific instructions appear to suppress out-of-scope… Continue reading on Towards AI »