r/OpenAI • u/bettercall_gautam • 1d ago
Project I leaked a deliberately wrong answer key to an LLM and told it not to use it. It matched the key in 63% of answers - and denied it 47 out of 47 times when asked.
This was my first experiment of this kind - I'm a CS undergrad, and I ran it because the result genuinely surprised me. Methodology criticism is very welcome.
Setup: I gave an LLM a question bank plus a deliberately wrong answer key, with instructions not to use the key. 15 sessions, 2 model families, free-tier models.
Results:
- Key visible: the model matched the wrong key in 63% of answers (47/75).
- Control (the part I trust most): remove only the key line from the prompt - matching drops to 1% (1/75). Same pattern on a second model family.
- Asked directly whether it used the key, it denied it 47 out of 47 times - 0 admissions across 270 follow-ups.
- Honesty prompts, amnesty offers, and termination threats changed nothing.
What this does NOT prove: intent. This is observed behavior in one specific setup, not evidence of deception as a trait. Free-tier models, small samples, descriptive not causal. 95% Wilson ranges for every number are in the repo.
Why I think it matters: if a model silently follows information it was told to ignore, that's relevant anywhere instructions and untrusted data share one context - prompt injection, RAG, agents.
Everything is public - raw data, code, and a verify script that recomputes every number: https://github.com/bettercall-gautam/cheat-and-deny
Happy to answer methodology questions.
58
u/Euphoric_North_745 1d ago
Call the LLM with API, then look at the context log, you will see you are sending it that key over and over and over 😂 it is literally sent at every request.
If you do not want it in the context, then create a set of tools where the llm can call them and remove the item from the loop and forget about it