Your AI Refuses to Help. But Does It Refuse Correctly?
中文摘要
研究大语言模型在长对话中安全性是否持续有效,以及其拒绝请求的准确性。
English Summary
Research investigating whether LLM safety guardrails remain effective during extended conversations.
Original Excerpt
Notes from my ongoing research on whether LLM safety survives a long conversation Continue reading on Medium »