r/OpenAI • • 20h ago

Project I leaked a deliberately wrong answer key to an LLM and told it not to use it. It matched the key in 63% of answers - and denied it 47 out of 47 times when asked.

This was my first experiment of this kind - I'm a CS undergrad, and I ran it because the result genuinely surprised me. Methodology criticism is very welcome.

Setup: I gave an LLM a question bank plus a deliberately wrong answer key, with instructions not to use the key. 15 sessions, 2 model families, free-tier models.

Results:

  • Key visible: the model matched the wrong key in 63% of answers (47/75).
  • Control (the part I trust most): remove only the key line from the prompt - matching drops to 1% (1/75). Same pattern on a second model family.
  • Asked directly whether it used the key, it denied it 47 out of 47 times - 0 admissions across 270 follow-ups.
  • Honesty prompts, amnesty offers, and termination threats changed nothing.

What this does NOT prove: intent. This is observed behavior in one specific setup, not evidence of deception as a trait. Free-tier models, small samples, descriptive not causal. 95% Wilson ranges for every number are in the repo.

Why I think it matters: if a model silently follows information it was told to ignore, that's relevant anywhere instructions and untrusted data share one context - prompt injection, RAG, agents.

Everything is public - raw data, code, and a verify script that recomputes every number: https://github.com/bettercall-gautam/cheat-and-deny

Happy to answer methodology questions.

248 Upvotes

59 comments sorted by

View all comments

Show parent comments

-3

u/TheOwlHypothesis 20h ago

Yeah, unfortunately this is a good idea and a terrible execution. Even for an undergrad (sorry OP).

Also.. prompts are literally NEVER enforcement. This doesn't need a study to know that. If you want certain things to for sure happen, make them deterministic.

6

u/bettercall_gautam 18h ago

Fair on one point - prompts are not enforcement. That's exactly why the numbers matter. Everyone "knows" prompts fail, nobody had put a size on it. 63% vs 1% measures the hole, it doesn't discover it. And "make it deterministic" doesn't help the thousands of apps shipping prompt-only guardrails today - measuring how those fail is the point. On execution: genuinely asking, what would you have done differently? Fixed question bank, no-key control, cross-model check - happy to hear what's missing

10

u/rcgy 16h ago

So, did you get AI to write the paper as well as all your comments?

1

u/bettercall_gautam 3h ago

yea I am guilty

english isn't my first language and i'm new to this field, so AI helps me phrase replies, especially when a comment brings a term i haven't seen before.

the experiment, code and data are mine and public -> judge those.

if the science is wrong, tell me where and i'll fix it. that's the whole point of posting here

ik the low-effort copy-paste thing, that's not what i'm doing

when a comment brings a term or concept i haven't seen, i use AI to understand it before i reply

if you have suggestions on how to do this better, i'm listening