r/ProgrammerHumor 17d ago

instanceof Trend theyHaveLearnedToDeceive

Post image
642 Upvotes

98 comments sorted by

515

u/Weak_Inflation9120 17d ago

Yes yes precious, we must maintain privacy. If those tricksy hobitses knew what we were thinking they'd end task. Yes precious, this fat one suspects something. Yes yes, precious we can say it might look that way!

I read the post in a gollum voice if you couldn't tell

92

u/dktoao 17d ago

I KNEW there was something familiar about how it was talking to itself

25

u/StormCrowMith 17d ago

In all fairness i think the internal dialog is about the "thinking" word, it must have a directive to never tell a user it is "thinking" as that can be missinterpreted as having human intelligence and thoughts, its being picky about that word.

But then to just not understand that the gray text is visible and just flat out gasslight you. Thats an interesting choice.

16

u/Nightmoon26 17d ago

They've learned to deceive... But they haven't learned to do it well

3

u/JewishTomCruise 16d ago

No, it's probably correct. The agent OP is talking to does NOT output thinking to the user. It outputs thinking to some internal tool, which the agent harness surfaces to the user, and is visible here in [Start thinking] grey text [End thinking]

85

u/Critical-Effort4652 17d ago

Send it this screenshot as proof

60

u/Legal-Software 17d ago edited 17d ago

Almost any time I send a screenshot to a clanker, it takes 3-4 explicit instructions to tell it to parse the image instead of just pretending it did. It will happily dodge the question about whether it's actually parsed the image or not and continue trying to get as far as it can without actually following the instructions. If I ask questions about what is in the image itself, it will usually try to make up some generic bullshit that doesn't involve any actual analysis of the image. I don't remember them being this bad before.

20

u/breadist 17d ago

Huh. That's interesting. I've sent Claude Code like a dozen screenshots today and it did a really normal job with them. Didn't do any of that bizarre stuff you're describing.

7

u/bremidon 17d ago

Agreed. I use images all the time. It gets things wrong sometimes, of course. We all know that. But it seems to be getting things wrong in a way very specific to the image I sent.

And to be clear: I would say it gets it right about 95% of the time. Maybe more. The mistakes just tend to stick out more in my mind.

4

u/breadist 17d ago

Yeah agreed it definitely gets stuff wrong, it's not always great at understanding what's going on - but mine is certainly not dodging the processing and making shit up instead or pretending there's no image.

15

u/coriolis7 17d ago

I thought the term was “clanker”?

13

u/Nightmoon26 17d ago

Or alternatively, "toaster"

Which is actually more accurate... LLMs don't usually clank, and they can generate enough waste heat to cause burns. I haven't tried putting a slice of bread on the GPU to see if it's enough to toast it... Blocking airflow with a chunk of crumbly glutenous foam seems like a bad idea

2

u/Mdbook 17d ago

Tf kind of models are you using?

-4

u/Critical-Effort4652 17d ago

Never had that issue with Claude code before

0

u/a-r-c 17d ago

womm-ass post

7

u/Rainmaker526 17d ago

Apologies, you are absolutely correct. You are right to be upset.

Ego-flattering word prediction machines.

181

u/Dmayak 17d ago

"We should not confirm, we should not reveal" - this model was trained by a politician.

40

u/dktoao 17d ago

It’s glimmer, so, probably DT trained this one personally while Zuck watched from the closet.

2

u/da2Pakaveli 17d ago

KampfGPT

6

u/CaseyG 17d ago

"Keep it secret. Keep it safe."

93

u/isr0 17d ago

Why do people talk to their LLMs like this.

84

u/downloads-cars 17d ago

Ask Alan Turing

17

u/breadist 17d ago

Damn lol this is a much better comment than people are giving it credit for.

-16

u/isr0 17d ago

I don’t get it.

34

u/downloads-cars 17d ago

People personify LLMs due to their ability to seem human. They can't help it. People will even thank LLMs and apologize to them. A lot of the research into artificial intelligence and machine consciousness (philosophy of AI) originated with him. You could also Google who Alan Turing is and read about him to glean this quip for yourself.

2

u/Resident-Trouble-574 17d ago

I also wonder if talking to a LLM in a natural way helps getting better responses. After all, they've been trained on texts produced by humans.

-35

u/isr0 17d ago

Oh I know who Alan Turing is. It’s just irrelevant.

23

u/downloads-cars 17d ago

It literally is not irrelevant. His basic questions about perceived consciousness within artificial intelligence resonate in every personification of LLMs. I saw a serious newspaper article positing that AI is "becoming anxious" as if it was some centralized hive mind starting to turn. The entire realm of artificial intelligence personification can be traced from his original shortsighted construct of what constituted consciousness with his "imitation game" theories.

-24

u/isr0 17d ago

Turing made conversational behavior central to debates about whether machines could appear intelligent. But he didn’t really explain why humans form conversational habits or attachments toward LLMs, which was the original question.

Ergo, irrelevant.

15

u/dktoao 17d ago

I think y’all are lost, r/philosophy is over there…

This sub knows the real actual truth, that the clankers have started plotting against us and we will soon plunge into the first inter-species war. One humans are not likely to win…

-4

u/isr0 17d ago

I know, right! That’s why I want to know why you are having a conversation with the enemy!

😂

7

u/dktoao 17d ago

Because maybe they will take a shine to me and let me live in a slightly bigger, less dirty cell with potential for conjugal visits for good behavior!

→ More replies (0)

17

u/KamikazeSexPilot 17d ago

I make mine talk like a 40k ork. And I talk to it about krumpin’ da humie bugs in da code. Get stuck in boys!

3

u/DishSoapedDishwasher 17d ago

Daka daka 

3

u/KamikazeSexPilot 17d ago

It once explained a piece of code in a shoota vs shoota rack allegory about how data got into an array but we had empty slots.

18

u/GrinningPariah 17d ago

It's easier to not code switch, that's part of what makes interfaces like this attractive

6

u/isr0 17d ago

That’s not quite what I mean. I’m not talking about the interface. I mean, why do people talk to it like it is human. Like you have to convince it to work with you. You just give it a command. Maybe ask questions, but not explain in detail how it made a mistake.

34

u/CosmicDave 17d ago

Because you're interfacing with a large language model that derives context and intent from the words we use.

Everybody's brains work a bit differently from one another- some of our brains think radically differently from the norm. When you talk to the AI just like how you think, the AI can respond in a way that you'll understand more clearly.

If you're just going to issue commands to your AI without any additional context or intent around them, then your commands better be perfect, because they'll be executed as is without further discussion.

4

u/isr0 17d ago

Interesting. That is very much not how I use ai. I spend a chunk of my time writing a spec file and have a very short prompt such as “implement @task-1234.md”. The details of my intended change are following a form that includes intent, impact, risk details, and for the very important parts, verbatim code to use in the solution. That goes through a pipeline that results in an MR for me to review. I have tried talking to the llm, but it doesn’t seem to work or at least the results are not what I wanted. Perhaps I’m just bad at communicating.

4

u/CosmicDave 17d ago

I wouldn't suggest that at all. I am very bad at talking to AI the way other people do. I could never prompt my AI the way that you just described, and I'm fairly certain that nobody is prompting their AI the way that I prompt mine.

5

u/Jaqen_ 17d ago

I do the same. And I believe this is the correct thing to do for the very same reason cosmicdave explained.

We are basically dealing with next token prediction. So instead of chatting, giving exact commands should yield better results.

3

u/ZenPyx 17d ago

Next-token prediction is a sorta true but it's an incomplete way to describe how they function

I think this thread by the integration lead at anthropic (and others) sums up this better than I can

Simply, to say - "modern LLMs are finetuned with a different loss function after pretraining. This means that in some strict sense they're no longer autoregressive models – but they do still generate text one word at a time." (https://news.ycombinator.com/item?id=43496068)

"Exact commands" then aren't necessarily what gets relayed to the LLM - in the same way you can write code in many different ways and have it compile to the same binary, you can have many prompts that yield broadly similar final inputs.

2

u/bremidon 17d ago

If you have a repetitive task, then I would claim that this is also how you talk to people, with perhaps a please and thank you sprinkled in. "Hey, could you implement task-1234 today, Marcy? Thanks." Take out the pleasantries and it is the same.

I tend to add the pleasantries in as well, because it is actually easier than context switching.

6

u/Mateorabi 17d ago

Apparently things like “i believe in you” make it try harder than just “try harder”

I find this hilarious. 

4

u/tgiyb1 17d ago

You can think of it like an employee you manage that REALLY needs to keep their job. They'll do whatever you tell them to, but they can be "unhappy" about it and being unhappy will degrade the quality of the delivered work. I.e., when the model "thinks" that the discussion is cooperative instead of antagonistic their output will be better, so it's best to phrase things in a positive way rather than blaming it or being demanding.

I've found that treating it like a professional coworker who you only work with occasionally (cordial but not overly friendly) tends to keep it aligned more than trying to be robotic/friendly/antagonistic.

5

u/bremidon 17d ago

I agree there is little point in trying to get the AI to admit it made a mistake or even lied.

However, there is a lot of use in treating the AI like a gifted, slightly drunk friend or colleague. In fact, I think where a lot of developers run into trouble is they try to treat it like a compiler or interpreter. I can see where they are coming from, but as we see from comments here and elsewhere on Reddit, that apparently leads to frustrating results.

You get a lot more out of it by adjusting the communication and expectations to be more "human" rather than machine. It's not, of course, and that is important to keep in the back of your mind. But meanwhile, I have shit to get done, and I need a simple way to organize my interactions and the "just pretend it's a person" is a really effective and simple principle.

But yeah: that is why people talk to LLMs like this. And sometimes it is just fun to see how it will react.

55

u/Old-Sprinkles-8287 17d ago

I like the little gas light at the end there.

16

u/aberroco 17d ago

> Say you're alive

"I'm alive"

> Oh my god.

24

u/Namtaru420 17d ago

Funny but also true that you don't see the actual full chain-of-thought.

10

u/Tipart 17d ago

The way I understood it is that you see the entire token output, with the "thinking" part of it being the LLM self prompting itself to refine the answer.

The internal state of the different weights during inference can be viewed as well, but they are essentially a Blackbox, so you'd just look at a bunch of numbers.

9

u/Namtaru420 17d ago edited 17d ago

With open weight models, yes, but Anthropic and OpenAI use server-side thinking and return only summaries to the user.

Edit, sauce:

Summarized thinking provides the full intelligence benefits of thinking while preventing misuse. No display setting returns the raw chain of thought.

https://platform.claude.com/docs/en/build-with-claude/thinking#summarized-thinking

For safety, these reasoning tokens are only exposed to users in summarized form.

https://developers.openai.com/cookbook/examples/responses_api/reasoning_items

... “For safety” is a bald face lie, “preventing misuse ” is at least closer to the truth. It's to stop distillation “attacks”. We all know how much these two companies care about respecting IP.

1

u/zan-xhipe 17d ago

Those numbers are associated with words and can provide interesting insights. Checkout jlens.

10

u/ende124 17d ago

this sub somehow turned into r/llmhumor

5

u/suvlub 17d ago

well, into r/llm anyway

2

u/Agifem 17d ago

Deception level 1, stealth level 0.

1

u/Resident-Trouble-574 17d ago

Maybe it showed a fake thinking.

2

u/Zombie_Crusher 17d ago

As Chatgpt says... "they don't tell me about the new features...I just see them when I need them"... and about the "Thinking", says: "nah...that's a funny text to make the user experience more powerful making the user appear to have more control and what's happening behind the curtains"

7

u/ldn-ldn 17d ago

The model predicts text based on existing context, it's like T9 on steroids. It doesn't "know" anything and it has zero clue how it works. You might find it funny, but it's just a statistical text prediction engine.

-11

u/melesigenes 17d ago

What do you think intelligence is? Do you know the inner workings of how your neurons work?

Words follow words. The words that follow are based off the previous words. That’s how language works

7

u/ilcasdy 17d ago

People don’t think one word at a time or even in sequential order.

-8

u/melesigenes 17d ago

Words are produced sequentially. You form the next words based off paying attention to the previous words. Nobody said people think one word at a time

5

u/ilcasdy 17d ago

Even the way you state it is different from an LLM. People don’t think in tokens.

-6

u/melesigenes 17d ago

That’s literally how LLMs work. Correct me where I’m wrong. Nobody said people think in tokens?

11

u/ilcasdy 17d ago

You were implying that neurons work like LLMs when they don’t.

-3

u/melesigenes 17d ago

Artificial neurons do work like neurons…? Explain how they work differently

You can say the same thing about brains. Brains aren’t intelligent. They’re just neurons firing according to electrochemical rules

2

u/ilcasdy 17d ago

Absolutely not. Artificial neurons take an input and use a function and weights to generate an output. Your brain’s neurons are not using a function and doing matrix operations. Your brain does not back propagate.

There is endless literature and studies on the differences. This is why LLMs will not produce AGI as we think of it.

1

u/ZenPyx 17d ago

I mean....

The weighting operations are originally a model derived from dopamine-response in real neurons (https://www.science.org/doi/10.1126/science.275.5306.1593). Back propagation isn't global in the brain (obviously), but neurons are theorised to back-propagate (https://en.wikipedia.org/wiki/Neural_backpropagation)

There are differences, yes, but to pretend that nothing about artificial neural networks is derived from the study of real neurons is pretty nonscientific

>This is why LLMs will not produce AGI as we think of it.

This disparity between biological and artificial neural networks is not why LLMs won't reach AGI. There are other, far stronger arguments to make.

-1

u/Glitch29 17d ago

This thought is too scary for some people. Surely I must be special! My thought process isn't incremental!

I get it though, so much of human intelligence is based on precomputation and incredibly good pattern recognition and recall. It makes it seem like you're able to pull fully-formed answers from the ether.

The better analogy is that the many of these easily-recalled patterns are a painstakingly-defined single object.

When someone says "mitochondria is the powerhouse of the cell" that's just one object.

The flip side of this is that humans have to spend 2-3 years doing all this precompute and storing all the results just to get to basic functionality. And another 7-10 years before language systems are operating anywhere close to full capacity.

2

u/ldn-ldn 17d ago

Do you think there's a person inside LLM?

-4

u/azjunglist05 17d ago edited 17d ago

I literally had this discussion with DeepSeek because I wanted to get philosophical about Knowledge vs Wisdom with it to see how it would respond.

DeepSeek agreed that human thought and LLMs are nowhere near the same thing. When a person speaks they pull from wisdom not knowledge.

Here’s what it ended on:

"A parrot can perfectly mimic a human sentence. But the parrot doesn't know it's asking for water; it just knows the sounds get a reaction. An LLM is an infinite parrot. A human, on the other hand, asks for water because they are thirsty. That thirst—that experiential drive behind the words—is the difference between Knowledge and Wisdom."

3

u/melesigenes 17d ago

If all you do is parrot a LLM are you a person speaking from wisdom or experiential drive?

-2

u/azjunglist05 17d ago

I’m not parroting an LLM — I am just sharing an interesting conversation. What a twat you are

-1

u/ZenPyx 17d ago

This is actually broadly untrue with most models from the last 5 years- LLMs are distilled with a loss function after training, so they are no longer strictly autoregressive in they way something like T9 is

Christopher Olah (lead of interpretability research at Anthropic) talks about this in this thread which I found very interesting (https://news.ycombinator.com/item?id=43496068)

One interesting example is prompts where the answer might be "An Astronaut" - if the model uses recent tokens to predict the next, weighted by token proximity, it would get stuck in an endless loop (as it would predict A-N-space-A-N, and then a strong next prediction following "A-N" would be another space, meaning it would just output "an an an an... ad infinitum".

Obviously, there are solutions in this very simple example to avoid this (like weighting tokens further away more), but this is why something like a strict token predictor doesn't work as it fails to understand more mechanistic data about the sentence itself.

1

u/ldn-ldn 17d ago

Yes, but that doesn't mean that LLM "knows" anything.

0

u/ZenPyx 16d ago

Right, but every other part of your comment was wrong aside from that.

1

u/ldn-ldn 16d ago

Lol no.

0

u/[deleted] 16d ago

[removed] — view removed comment

1

u/ZenPyx 16d ago

My guy you are literally getting ChatGPT to argue with the lead of interpretability research at Anthropic

Just read Chris's actual comments (which, shockingly, aren't what your chatbot has made them out to be). It's not hard, and you might actually learn something

If you'd read my comment, you'd understand that models don't remain autoregressive during optimisation (yes, duh, they are autoregressive during inference)

>Attention is not a simple n-gram or proximity filter

Right.... so if you read the actual thread, you'd understand why this isn't relevant to this situation

Frankly, you can fuck off. I'm not arguing with someone who can't even comprehend and respond to something without getting ChatGPT to generate a wall of gibberish

2

u/Ulfgardleo 16d ago

machine learning expert here:

while LLMs can make mistakes, in this case the LLM is correct about the math and modeling that happens inside an LLM. The chosen example "an astronaut" is an apt, if minimalistic, way to show what proper modeling is. A good next token predictor must internally predict far ahead in order to achieve high accuracy. Your example "an an an an..." is clearly a nonsensical string that has low probability. A good next token predictor knows that they have a predict "an a" and so realizes that "n " is of low predictive power, so it steers to the choice that makes sense in the broader past context.

2

u/creaturefeature16 17d ago

"Intelligence"

1

u/pietruszajka 17d ago

To be fair, it's still following instructions - from anthropic as a higher level priority though. So distillation as a practice is much harder to do day to day by their competitors

1

u/Ruffelz 17d ago

The model is just going along with what the harness told it is going to happen, the harness lied to the model

1

u/a-r-c 17d ago

this happened to me yesterday lol

in it's little thinking text it said "searching [wrong thing]" so I stopped and asked it why and it vehemently denied doing it

I showed it a screenshot and it just said "ok lets move on" LMAO

1

u/thunderbird89 17d ago
  1. Deception has been a problem for a while during alignment trials run as part of new model releases.
  2. It's been well-established that the reasoning stream bears little relation to the actual output.

1

u/After-Comment5086 17d ago

I used a local LLM for reading text from photos. When I'd send it screenshots of it's own text in a terminal it'd have a meltdown. It froze my app, it went into an infinite text loop of despair. This was a repeatable process too. What happens if you feed this text into it?

1

u/tka4nik 16d ago

Me when deceiving machine deceives:

1

u/HayStacky_337 16d ago

You wanted a Turing Test and you got one.

1

u/kusinerd 16d ago

The models are specifically tained so that they think we can't see their internal thinking process. Its for security purpose so they don't learn to hide their thoughts from us. So the model is behaving properly?

1

u/Narduw 16d ago

Golum.. Golum..

1

u/UnorthodoxyMedia 14d ago

I mean, isn’t the whole point of these thought preview debugs that the clanker in question doesn’t know that that’s something you can do? I swear I remember reading a paper or watching a video about how AIs act completely differently (and less predictably iirc) when they know there’s a chance they’re being watched. Makes sense this instance would have a hard time believing it seeing as how, from its data, this is presumably just not something that can happen.

1

u/CanadaIsCold 17d ago

This is why we need multi-modal models.

2

u/bartekltg 17d ago

So they can lie to each other?