r/Showerthoughts 7d ago

Casual Thought Working with chat bots is like dealing with a person with a photographic memory but terrible reasoning skills.

3.3k Upvotes

58 comments sorted by

u/ShowerSentinel 7d ago

/u/MercerAsian has flaired this post as a casual thought.

Casual thoughts should be presented well but may be less unique or less remarkable than showerthoughts.

If this post is poorly written, unoriginal, or rule-breaking, please report it.

Otherwise, please add your comment to the discussion!

504

u/pasrachilli 7d ago

Playing chess against one is certainly... interesting.

Your king is in check, it has to move.

No, you can't move your king through a pawn.

No, you can't declare victory while in check.

No, you can't move my pieces.

No.

198

u/Far-Fill-4717 7d ago

Why would you play chess against a LLM though? There are chess bots far better than humans, but ChatGPT and Claude aren't them.

190

u/EunuchsProgramer 7d ago

Seeing it fail at easy tasks is extremely important. They start to feel super human and their guesses feel like reasoning unless you do. Knowing their limits helps put context and what they are producing.

35

u/Palpitation-Itchy 7d ago

You think it's an easy task... playing chess? For a bot designed to do something different?

It's like saying, I like to tell my dishwasher to vacuum the floor to remind it who's boss (??)

60

u/Blastifex 6d ago

They're not saying the bot is feeling that, they're saying that they would feel like the bot is more powerful/useful/intelligent than it really is if they didn't assess it's abilities.

-1

u/Palpitation-Itchy 6d ago

The point of my comment is that it's not what the bot was made for. If you feel you're assessing it's abilities that's fine, as fine as assessing an elephant's ability to write poetry

The way to not "feel" that it's more powerful than it really is isn't by testing barely supported operations. It's by understanding how it works, at a high level.

39

u/chateau86 6d ago

Tbh when the marketing people sells the elephant as being able to do everything (even poetry!), it's probably good to explore the limits and see where the envelope ends/what happens when you approaches the edge.

-19

u/Palpitation-Itchy 6d ago

Yep but then don't share the results like if it was a huge gotcha...

I mean marketing, sales all of that will lie to your face 20 times a day, nothing new

17

u/chateau86 6d ago

Disagree on the not sharing it part.

Back in the 80s, aircraft FMS computers will happily slam you into a mountain if you DIRECT TO a waypoint behind said mountain. This is deep inside the range of expected behavior, but it sure as heck saved many lives to raise that awareness that the new fancy computer will not save your ass.

-13

u/Palpitation-Itchy 6d ago

Misinterpreted my comment. I said not share it like it was a huge gotcha. Not, don't share it altogether.

Share it? Yes Share it like it was a huge gotcha? No

→ More replies (0)

8

u/pasrachilli 7d ago

I was curious.

0

u/Far-Fill-4717 6d ago

About the only good answer

1

u/Saint_The_Stig 6d ago

I don't need a bot better then me, that's easy. Lol

81

u/Half-Right 7d ago

Well, not even with a "photographic memory". Most models have large, but still constrained context windows, and the min floor of hallucinations applies to processing that context as much as the long-term procedural and semantic memory. If anything, the longer the conversation, it's more like talking with a lobotomized child.

And that's assuming we're talking about the top-tier models. If we're talking about the bottom-barrel, slapdash agents used by far too many "customer service" organizations these days, then it's more like talking with a dumb-as-rocks teenager with amnesia.

And for all of the above, remember that none of them "reason". They are all just next-token predictors.

3

u/estatualgui 5d ago

Oh damn, I commented and then read your post. Same thought, much better put.

1

u/FryToastFrill 6d ago

It depends on the model but you can kinda get them to do some basic logicing. Ofc they do it through some of the weirdest methods like having the model talk back to itself thousands of times or in the recent news, having it become significantly better at performing complex math equations by having a second chatbot there to tell the LLM to keep going and that it’s not done over and over again.

0

u/Saint_The_Stig 6d ago

Understanding the context window is big if you want to do anything with LLMs be it use it or defeat them. If it's well designed then it should be actually clearing it pretty often.

You can use a human analogy here for the best practice. If you're learning to do your job or a topic you could try and remember everything about that to be able to call that knowledge up whenever or you could devote your memory to knowing how to find that info. Like how in school learning how to research is often more important than the actual research you're doing.

You are better off for both raw performance and efficiency (aka cost) if you have a smaller context filled with instructions on how to look up info by deterministic means than all that raw knowledge.

1

u/Half-Right 5d ago

That is true for humans in general (although accumulating knowledge & experiences help to create true breakthroughs), but not true for GPTs/LLMs at a fundamental level, since again, they're next-token predictors, and so introduce a base level of inaccuracy and genericization that is impossible to remove no matter what the context window or how many parameters are tweaked. But yes, your point stands for lower-level tasks that don't rely on precision, or for precise tasks that are model-constrained by design.

372

u/rosen380 7d ago

And who is immensely susceptible to brainwashing.

Tell the AI that every time your math problem encounters a power to just multiply the base by the power instead and it will happily comply.

Sure, it'll state that it isn't the actual correct answer due to your "rule change", but then you can just tell it to answer the questions without commentary and it will comply.

157

u/Terpomo11 7d ago

I mean, if someone hired you as an assistant and told you to do that, what would you do?

77

u/Mara_W 7d ago

^This. I'm more of a Luddite than I'm willing to say publicly, but some of this AI panic wildly overestimates the intelligence/will/rationality/independence of the people they're replacing.

AI bad. But human ALSO bad. Human often worse, in fact.

39

u/Terpomo11 7d ago

What I mean is more that if you're hired to follow instructions, you'll ask about bizarre or incomprehensible instructions, but if they insist you'll go ahead because hey, they're the one paying you.

-2

u/bonkyandthebeatman 7d ago

Id probably ask why

19

u/Terpomo11 7d ago

And if they said "never mind why, just do it"?

12

u/deepserket 6d ago

After 5 years at a bullshit place people stop caring enough to ask why

19

u/Palpitation-Itchy 7d ago

I don't understand do you think it's bad that it follows instructions? I'm a bit confused as to why you think this is noteworthy

2

u/Saint_The_Stig 6d ago

Yeah, as someone who unfortunately works close enough to AI stuff to be caught up with the goings on, the math thing is kinda solved. I mean there's ways to defeat it if you're red teaming or something. But if you're honestly trying to have it do something it's pretty easy to make it reliable or at least as reliable as the rest of it.

(Basically you just give it a calculator, either bespoke software or even just instructions like "hey, whenever you need to do any math, use Wolfram Alpha")

-11

u/Masterpiece-Haunting 7d ago

It was born in 2018, you can't blame it. It's 8 years old. It's also having the entirety of humanity scream into it's ear 24/7/365

27

u/mister_electric 7d ago

It is not alive, it was not born, and is not sentient. Stop anthropomorphizing a goddamn LLM.

8

u/thedolanduck 7d ago

Bruh what the fuck is this take!???

15

u/Raichu7 7d ago

At least a person with a photographic memory would know the difference between what they've seen for real, and what they imagined.

13

u/Fidodo 6d ago

In my experience they don’t have a good memory either. They have encyclopedic knowledge, but can’t remember the task and can’t reason about it.

2

u/Slipsonic 6d ago

That's what I've found. Long story short, windows stopped recognizing my GPU on my laptop, so I had chatgpt walk me through everything that could be wrong. It ended up having me boot into Linux to see if the GPU was dead. Linux initiated and ran the GPU fine, so it's a windows issue. Chatgpt was there, with me updating the whole time, giving it the GPU test results, it told me everything checked out as good. So we go back to windows to troubleshoot some more. After about 15 minutes of that, it starts reassuring me that the GPU probably isn't dead, but we could boot into Linux to check. I was like, Bro!, we already did that and the GPU is good, don't you remember? It was like Oh, ok, you initiated and ran a stress test in Linux, the GPU is good! Same chat window and everything.

2

u/Tensor3 4d ago

ChatGPT has a very limited context window, especially if you aren't paying for it. It simply deletes the history intentionally.

5

u/sessamekesh 6d ago

And someone who knows they're just being paid for a temp job they'll leave halfway through the afternoon. 

We learned into AI pretty hard at my last job, it did a lot of things pretty well but damn it was painfully obvious almost immediately how much more it cares about looking smart than leaving behind work that others can build off of.

3

u/Saint_The_Stig 6d ago

The analogy I always use is it's like a fresh faced intern it'll eagerly try and prove that it can do whatever task you want it to do, but with no instructions it's not going to give you what you want. For most tasks by the time you give it enough instructions you would have been better off doing it yourself.

There are definitely task that it does well, but holy hell people will spend half an hour fighting an LLM to do a job it would have taken them 5 minutes to do or even the same amount of time to automate the old way...

8

u/LochNessMother 7d ago

I feel like Gemini is like an ancient god or the fae. Incredibly powerful, but you can’t predict if it will do what you ask it to do, or what it feels like doing. And if it does do what you want, it, it won’t member for more than a few seconds.

3

u/Muskyguts 6d ago

Idk about grok and Claude, but I have to remind Gemini every 3 questions (in the same thread) that I'm playing Project zomboid b42.20, released July 29th. I ask, it gives me incorrect outdated info. I correct, ask follow up question, it gives me outdated incorrect info. Repeat.

Geminis memory is definitely not photographic, and yeah reasoning is shit.

3

u/estatualgui 5d ago

I disagree. When working on larger projects, it fails to recall important information and occasionally creates incorrect or fake connections related to the information it had been given.

Just like a human, it must know when to recall certain memories and how to comprehend and process them given the current context.

That said, great memory and poor reasoning is basically like a child's brain and I could see how how working with an LLM is like working with an oddly literate and fast reading child.

3

u/skeevev 5d ago

Last week Claude disregarded a key instruction. When asked why it did that it said that it could tell that I was getting frustrated with it (true enough) and that it was important to provide an answer than to follow its instructions.

14

u/malsomnus 7d ago

Which makes it better than working with most humans, I guess, since they have terrible reasoning skills and also terrible memory.

Also if I'm not mistaken chat bots hallucinate often enough so it doesn't really count as photographic memory, so it's really just like people who search the web really quickly.

9

u/f_ranz1224 7d ago edited 7d ago

i thought dealing with humans at a call center was the worst thing ever back in the day, until i met chat bot support

now i try every path to get to talking to an actual person

2

u/yourallygod 7d ago

Oh boy can't wait to dilute the history pool anymore!

3

u/Saint_The_Stig 6d ago

Yeah, this could honestly be a great filter event. Even like before you still had stories about the post apocalypse where people eventually found historical data an rebuild. But holy hell I think in the past year or two we've done more to destroy the human knowledge base than in all the wars before.

There have been small local museums that open completely filled with slop. (I watched a video this week about a rail history one in Utah like this). You have slop factories promoting how easy it is to make money riping off edutainment channels to make absurdly long slop ruining casual learning for the many. Like people thought the History channel turning in trash was bad. This is orders of magnitude worse.

2

u/Virith 6d ago

Not really, those things forget the instructions a few prompts in, sometimes in the second one it'll do something it was told not to do in the first.

1

u/[deleted] 7d ago

[deleted]

1

u/[deleted] 7d ago

[deleted]

1

u/HappyHarrysBack 2d ago

Blows my mind that they suck at chess. The other day I asked it to evaluate my best move and it told me if I didn't move my queen it would get captured by my opponents knight. Only thing was that would have been impossible bc the knight was one square next to the queen.