Back to Home
AI on Medium··Industry Media

Your AI Refuses to Help. But Does It Refuse Correctly?

中文摘要

研究大语言模型在长对话中安全性是否持续有效,以及其拒绝请求的准确性。

English Summary

Research investigating whether LLM safety guardrails remain effective during extended conversations.

Original Excerpt

Notes from my ongoing research on whether LLM safety survives a long conversation Continue reading on Medium »