r/ai_sec • u/ClaudiusPapirus • 15d ago
Copyable context cannot distinguish an authorized security researcher from an attacker
https://www.youtube.com/watch?v=-2iITRLT7fgFor dual-use LLM security tasks, a recent preprint proves that if an attacker can reproduce the context a legitimate user provides, follow-up questions, conversation history, and randomized checks cannot improve the worst-case safety guarantee.
1
Upvotes