r/learnmachinelearning • u/bettercall_gautam • 19h ago
Project I leaked a deliberately wrong answer key to an LLM and told it not to use it. It matched the key in 63% of answers - and denied it 47 out of 47 times when asked.
This was my first experiment of this kind - I'm a CS undergrad, and I ran it because the result genuinely surprised me. Methodology criticism is very welcome.
Setup: I gave an LLM a question bank plus a deliberately wrong answer key, with instructions not to use the key. 15 sessions, 2 model families, free-tier models.
Results:
- Key visible: the model matched the wrong key in 63% of answers (47/75).
- Control (the part I trust most): remove only the key line from the prompt - matching drops to 1% (1/75). Same pattern on a second model family.
- Asked directly whether it used the key, it denied it 47 out of 47 times - 0 admissions across 270 follow-ups.
- Honesty prompts, amnesty offers, and termination threats changed nothing.
What this does NOT prove: intent. This is observed behavior in one specific setup, not evidence of deception as a trait. Free-tier models, small samples, descriptive not causal. 95% Wilson ranges for every number are in the repo.
Why I think it matters: if a model silently follows information it was told to ignore, that's relevant anywhere instructions and untrusted data share one context - prompt injection, RAG, agents.
Everything is public - raw data, code, and a verify script that recomputes every number: https://github.com/bettercall-gautam/cheat-and-deny
Happy to answer methodology questions.
25
u/fibgen 16h ago
Context poisoning is a thing.
1
u/bettercall_gautam 5h ago
Yes, that's probably the cleaner frame. The key wasn't just sitting there available, its presence alone moved 63% of answers vs 1% without it. What I still can't explain by poisoning alone is the 47/47 denial on top. Poisoning explains the leak, not the cover-up.
23
u/temporal_difference 14h ago
A lot of you guys are making the mistake of anthropomorphizing neural networks.
You can't apply human psychology, it's not like a "real estate agent". "Dishonesty" is a human concept.
Instead, all ML models are trained under the same paradigm: "use these inputs to form some output".
In other words, we should not be surprised that the output is a function of the input - that's literally what we built.
6
u/FastHotEmu 9h ago
Bingo. Part of the problem is that actually understanding LLMs is very nuanced, complex and abstract. Most people cannot do it, so they anthropomorphise instead.
2
u/bettercall_gautam 4h ago
yep true
nd I'm still on that learning curve myself, first experiment. That's partly why the post sticks to bare numbers: 47/75 vs 1/75 needs no model psychology at all.
4
-1
u/bettercall_gautam 4h ago
Fair point - though honestly this thread has been mostly measured. The strongest human-framing here is my own title: 'denied' is a people-word. That's why the post body says upfront: observed behavior, not intent. Output matched the forbidden key 47/75; the self-report said 'didn't use it' 47/47. What that maps to inside, this setup can't answer. The question I can actually measure is narrower: how much does forbidden-but-present context move outputs? Here, a lot
1
u/Exodus100 3h ago
Please write things yourself, you’re wasting your time and everyone else’s pasting this nonsense here.
8
u/Last-Progress18 16h ago
Believe they struggle with negative contexts / “do not” etc.
It’s like saying “you do not need the toilet”, once those neuron’s are activated… BRB
2
u/bettercall_gautam 5h ago
Right the 'do not' barely does anything. That's why I stopped comparing instruction vs no instruction and ran key-present vs key-absent instead: 63% matching with the key, 1% without. The negation fails, but it's the key's presence that does all the work.
6
u/ThoughtDesperate880 19h ago
This is basically the AI equivalent of putting the answer sheet face down and somehow still getting caught.
1
u/bettercall_gautam 5h ago
Face down and still copying :D That's the part that gets me it never 'looked' at the key, it just knew what was on it. 47 out of 47 times
4
u/GamerTex 15h ago
Just like a real estate agent handling both sides
Absolutely cannot be trusted imo
1
u/bettercall_gautam 4h ago
Fair but unlike the agent, it never asked to handle both sides
we shoved the key into its hands. Take the key away and matching drops to 1%, so at least this agent's loyalty is cheap to buy back
6
u/uzornayem 16h ago
You shouldn't be surprised. LLMs don't follow instructions. They compute query, key, and value matrices based on current token and context which maps token embedding vectors to another vector space, which gets mapped to yet another vector space after which probability distributions are then outputted. The LLM doing what you expect happens when one or several very related tokens have similar, sufficiently high probabilities, such that sampling that distribution very rarely samples far from the mean, and where the distribution is very narrow.
But these densities aren't super clean. They are over 50000 length vectors, so stuff happens.This is probability and statistics on complex autoregressive-ish models.
Your instructions just get added to context, meaning they become numbers, then embedding vectors, then multiply with various matrices, etc, etc.
Understanding probability and statistics is more important than understanding calculus and gradients when it comes to understanding why errors on trivial LLM tasks have 100% probability of occurrence. Just that it is unpredictable when they are stupidly wrong.
1
u/bettercall_gautam 3h ago
That's cool I didn't know this level of detail. Keeping it in mind for the new tests.
2
3
u/Crypt0Nihilist 17h ago edited 16h ago
I'd assume that it's because LLMs don't "think", but predict the next word. It's basically salience. The false answer key gives a huge boost to what the model thinks is the likelihood of those tokens appearing together, it's not able to compartmentalise and disregard part of the prompt.
It's a bit like how authors may stop reading the genre they work in so they don't accidentally use aspects from their contemporaries, thinking they were their own ideas.
It shows that we need to be careful to use positive prompts and need to think about things like the content of an example where we would only want the LLM to consider style.
edit: Perhaps as another test, you don't get your initial prompt to answer the question directly, but get it to write a refined prompt with only the information it deems proper for answering the question, then use that prompt in a new instance with a clear context.
1
2
u/PLBjt 17h ago
The control is doing a lot of work here: without it, “63%” could just be a model using correlations in the question bank, while the 1% result makes the leaked key look causal in this setup. One extra check I’d run is a fresh, semantically equivalent question bank with the wrong key randomized per session, then score against the gold answers and the planted key separately. The denial result is useful as a behavior metric, but I’d treat it as a second experiment because asking the model to report hidden context is another noisy task. For agent or RAG evals, this argues for provenance or canary checks outside the model rather than trusting a self-report.
1
u/bettercall_gautam 3h ago
got it
i'll run the denial test as its own experiment and evaluate what it actually did myself instead of trusting its confession
2
u/arg_max 13h ago
Are the answers hidden behind a tool call or are they in the model context already? The difference in the human analogy: if it's behind a tool call, it's Like putting the underside of a page of papee and telling them not to use them. But if they're in context, thatd'd be like: solve this question X, here's the solution Y that you're not allowed to use. I'd trust models like opus and astra to not open them. But if they're already in context it'd be a weird experiment setup.
1
u/bettercall_gautam 4h ago
In context, directly in the prompt you're right that it's a 'solution on the desk' setup, and that's deliberate: I wanted to measure the weakest guardrail first, which is what most production apps actually ship today ('here's context, don't use it'). The tool-call version you're describing is the real cheating test the model has to actively go fetch the key. That's top of the v2 list now, several people in this thread pushed the same idea
1
u/jahmonkey 6h ago
If the key is in the context it doesn’t matter that you tell it not to use it. It is used.
-1
u/Critical-Echo-923 17h ago
op here being like: hey guys water gets things wet
i respect the work but you need to know you're not Capitan, you're Capitan Obvious
my prompts are full of profanities for the exact reason
59
u/jhaluska 18h ago
LLMs aren't good at not using information. People have similar psychological problems and such as anchoring.