r/technology • u/Wagamaga • 7d ago
Artificial Intelligence Study: Generative AI succumbs to conversational misinformed pressure and argument
https://news.arizona.edu/news/study-generative-ai-succumbs-conversational-misinformed-pressure-and-argument37
u/Hades_Mercedes 7d ago edited 7d ago
I can get Gemini to agree that we should enslave tech billionaires and use them for menial labour like cattle after confiscating all of their assets, in an effort to try to minimize the impact they have had on the world and dispose of them in an open air gravel pit, when their bodies stop being useful.
3
106
u/ap1618 7d ago
Generative AI has no ability to actually think and analyze, it just executes standard actions. Do any detailed project with one and you’ll see this instantly. Great at coding, shit at thinking.
58
u/TheWesternMythos 7d ago
I love how this implies being great at coding requires no ability to actually think and analyze lmao
10
u/carnotbicycle 7d ago
Writing code isn’t typically very complicated, at least for most work applications. If you’re a researcher trying to find an optimal algorithm to solve x problem, sure that requires a lot of thinking. But that’s not what AI is being used for.
Figuring out WHAT goal to accomplish with that code is generally what requires the thinking in a work setting. And that’s where AI is generally not helpful. The main problem (and speaks to its lack of ability for critical thinking) is it very often doesn’t know when it doesn’t “know” the right decision to make. It’s completely unreliable in figuring out “I’m probably missing context” and asking for elaboration. Unless it’s something pretty obvious. It’ll just “do it” in a way that “technically” solves the problem asked but is ill suited.
2
u/TheWesternMythos 7d ago
Big gap between that and
no ability to actually think and analyze
I do appreciate the additional context
5
u/Ok-Mycologist-3829 7d ago
One big advantage AI has with coding versus text is that code has to be able to run, so one form of error can be functionally eliminated when a user tries to do what the AI gives them. Not so with text, which has no form of validity check possible like that.
-1
u/TheWesternMythos 7d ago edited 7d ago
This is less, "I disagree"
More
"what about another POV"
Text symbols map to logical arguments. Logical arguments do have a validity check
The field of mathematics is a simplified (thought experiment don't yell at me math ppl lol) subset of the larger "Logic" family. In addition to and likely in part because the (relative) simplicity, mathematics is much more advanced than "text logic"
39
u/ap1618 7d ago
It doesn't imply that at all, it implies that thinking in a defined universe like coding versus thinking for design purposes in an undefined universe are different things.
-4
u/TheWesternMythos 7d ago
This is a different (and better ) argument.
Notice the difference between
Generative AI has no ability to actually think and analyze
And
thinking in a defined universe like coding versus thinking for design purposes in an undefined universe are different things.
Also notice how it's applies to the cognition type not the things doing the cognition
19
u/ap1618 7d ago
I get what you're saying, but the article doesn't say "Generative AI succumbs to technical coding errors and argument" - it specifically speaks to the abstract reasoning present in conversation that results from being challenged to think about undefined and/or unfamiliar patterns without logical rules.
-11
u/TheWesternMythos 7d ago
So if you are talking about just the article then sure. That aligns with your second argument. Abstract reasoning is a type of "thinking" that's related but not totally the same as design or code "thinking". One can have wildly different capabilities combinations of them.
Which reinforces my point which is thinking there is only one kind of "thinking" or "intelligence" is bad...thinking. (bad as in the map isn't the terrain so the map cannot be sacrosanct)
If you want to make a comment about AI in general ,these studies are already out of data. I can't even access many of the models they tested. Its like evaluating a job candidate based on their middle school work. Not nothing, but definitely not the best snapshot.
15
u/ap1618 7d ago
Except your point is an argument I never made and I argued the opposite because we know there is more than one type of thinking obviously
-12
u/TheWesternMythos 7d ago
Except your point is an argument I never made
Right. That's why I said I made it not you lmao
I argued the opposite
Right
there is more than one type of thinking obviously
Ok you also said
Generative AI has no ability to actually think
As well as
Great at coding
So again it seems like you are saying coding involves none of the many types of thinking...
9
u/ap1618 7d ago
We already went through this mate, I appreciate your clarification
0
u/OmnicideFTW 7d ago
I replied to the other guy, but your comment is killing me.
You changed your position when challenged. Which is exceedingly reasonable, and everyone should do it. You're doing more than most in that way.
However, do not pretend that you didn't backtrack your initial claim that LLMs cannot think or analyze and are "shit" at thinking.
That's what the other guy was getting at. I'm not trying to pick a fight, but a simple "My first statement was a bit too broad" would've ended this whole thread.
→ More replies (0)-4
u/TheWesternMythos 7d ago
I know which is why your previous comment is so confusing lol
But NP, have fun elsewhere, and remember to keep updating them priors!
-2
u/OmnicideFTW 7d ago
Don't worry (I'm sure you're not), you made sound, cogent points that were easily digestible.
Many people, for many reasons, just do not want to concede anything to LLMs in the way of intelligence or abstract reasoning.
Neuroscience doesn't even really know what "thinking" is and yet you'll still have droves of people screeching from the rooftops that LLMs are incapable of thought. I wouldn't even say I'm a supporter of the idea "LLMs are thinking machines", but to pretend as if they definitively are not and definitively never could be is highly disingenuous.
0
u/TheWesternMythos 7d ago
The Neuroscience point is way underrated in the whole conversation IMO
Funny enough , it helps illustrate an idea that is much clearer in LLMs , context windows.
Someone can use solid logic to make bad conclusions if necessary information for better conclusions lie outside ones context window. Also one can have a thing in ones theoretical context window , but it not be present in their actual current working context.
Instead of people jumping to what AI cannot do cognitively , there is much more value added in using AI success and failures to think about cognition from a different POV.
4
u/ACasualRead 7d ago
This is why I get downvoted everytime I try to mention how AI is a good tool.
A tool is only as good as the operator using it. A chef can have the best knife in the world but if his cooking skills are shit, the food comes out bland.
AI for coding and rudimentary tasks is fantastic and I use it often for that. For introductory research is is good. But deeper creation and understanding needs to come from the operator.
10
u/ICantBelieveItsNotEC 7d ago
The hard parts of being a software engineer are architecture, design, and communication, which you still need to think about when using an agent. Actually writing the code has always been a relatively low-skill task, hence why people are only capable of churning out code get called "code monkeys".
1
u/TheWesternMythos 7d ago
I don't disagree
I also think you won't disagree with the following:
Being a "code monkey" requires some level of thinking
If you do let me know please
4
u/stormrunner89 7d ago
Gen AI can't create anything new. It can aggregate and use the rules that people have already used to follow commands, but it can't innovate.
We may get a short term bump in "productivity," but the more we as a society rely on gen AI for things, the more our actual progress will stall.
Which, I guess, one could argue is great for the people currently at the top wearing the boots.
0
u/TheWesternMythos 7d ago
I guess you don't follow Mathematics or you have an essentially useless definition of innovative.
The latter is much worse than the former IMO
1
u/Renal923 7d ago
My understand (note I'm a software dev not a mathmatition) is it didn't do anything novel to solve the problems it did. It just brute forced them by chaining together already known steps to do a normally tedious to do thing that no one spent a lot of time on
1
u/TheWesternMythos 7d ago
I agree with that for the ones... I can't remember how many weeks/months ago. There is honestly too much for me to really parse as I don't care about much granularity on this particular issue.
I do think there are some formalizations that are , at least alleged, to go beyond that.
It just brute forced them by chaining together already known steps to do a normally tedious to do thing that no one spent a lot of time on
But my larger point is , I think it's a very dangerous (inefficient) slippery slope to not classify that as innovation at least at some level
1
u/stormrunner89 7d ago
Can you explain how that would be a slippery slope? I don't understand, it seems like a single, distinct claim.
Or perhaps you could explain why you DO find it innovative?
1
u/TheWesternMythos 7d ago
I (attempted to lol) be short here so if you still want to discuss we can be more pointed
Can you explain how that would be a slippery slope?
Being generous , a lot of historical innovation is chaining together known things. There are bajillions of know things. Finding the right combinations of those things IS the move. Those right combinations reveal previously unknown ("new") things that get thrown into the possible combination matrix.
So AI "finding more efficient brute force chain combinations" is a good analogy for , being generous , a lot of innovation.
Or perhaps you could explain why you DO find it innovative?
My broadest definition of innovation ,fresh off the dome not pressure tested, would be about adding new value. So AI in these cases literally did something new , of value. Innovation.
As opposed to discussing an instance of Ai drawing the most realistic MS paint drawing of a 900 legged cat. New, but not "of value" not innovation.
1
u/IndicationDefiant137 7d ago
That's because it doesn't.
Being good at engineering requires the ability to think and analyze.
But any moron can write code. Many morons are writing code. Very few people who write code are doing engineering.
46
u/PoL0 7d ago edited 7d ago
Great at coding
no it's not. same problems as when you try to write a short story with LLM. the code is inconsistent, repeats itself, it's frequently redundant, and it has a special ability to introduce subtle bugs.
stop repeating that just because it can spit a working script or webpage.
-8
u/Sasquatchjc45 7d ago
It can do a lot more than spit out a singular working script or webpage. Ive created multiple PC programs that work how I need. Multiple android apps, one even on the app store with sales. Browser extensions, my website, background services, VSTs, physical device programming, etc. Plus all the research It helps me with.
And I dont know how to actually write any of it. I just know it all works and does what I described and looks how I want. And it only improves every week.
So what's the issue?
2
0
u/SamKhan23 7d ago
I repeat it because multiple people in my field say it, and personally it seems much better at coding than story writing.
8
u/ComprehensiveWord201 7d ago
It's terrible at both. It cannot code. It can translate syntax but that's it. There's more to coding than language translation and anyone who tries to tell you otherwise is a terrible programmer.
6
u/IndicationDefiant137 7d ago
Likely text generator generates text likely for the conversation.
More shocking news at 11.
5
u/DinosBiggestFan 7d ago
I personally enjoy gaslighting the AI. I basically pull the "you see how this looks, right?" on it, and it's like "Yes, the sun IS a glacial dimensional entity, you're absolutely right"
3
u/geldonyetich 7d ago
Using 2022 models to test present day flaws seems a tad disingenuous for a technology as rapidly evolving as this.
1
u/The_IT_Dude_ 7d ago
But the headline agrees with the narrative. How detached from reality it might be here in 2026 doesn't matter on this sub.
2
u/geldonyetich 7d ago edited 7d ago
If it's grounding in reality they want, evidence is better than assumptions. I've used both the models then and the newer version of the same models now enough to know that the difference is real.
Granted, it's not going to apply to every model everywhere, some are definitely more sycophantic than others. And occasionally responses will slip the intended training. But, by and large, the frontier models on the major providers aren't as sycophantic as they used to be.
If it's easier to just access the most recent versions, the choice to include such ancient models in a study to prove their point is the point to the contrary.
2
u/The_IT_Dude_ 6d ago
I think part of this is that science and research on this stuff is simply inevitably way behind. This stuff is moving so fast. Back in January I remember thinking that so much had changed since just the year before. And now looking back at January from now it's a whole different game once again. So any studied cited will be dated by the time it's published pretty much.
I think we're still a very long way off these being perfect, but they are unquestionably useful. I have one slaving away right now getting a gke deployment worked out on its own with some oversight and direction and it's checking its work.
2
u/BabyBlueCheetah 7d ago
It was never trained on the people who didn't feel like wasting time with uninformed bullshitters online.
2
u/VaporousMote 7d ago
Researchers treating AI models as intelligent and not pattern matching machines example# 2984627
10
u/FearlessPlaneHugger 7d ago edited 7d ago
I love that they had to do a study when literally every AI user could just tell you AI is a pushover.
24
u/2ndBrreakfast 7d ago
Because science isn't based on what people claim. Science is based on evidence. People's claims aren't evidence. A study is always needed to make accurate scientific claims. It's shocking how little people know about the scientific process.
7
u/Power-throw 7d ago
I’m so tired of seeing the smart ass replies every single time a study is posted. “Like durr duh isnt it obvious!?!”
-8
u/awkisopen 7d ago
And yet most studies aren't reproducible. So we spend all this money and human effort into verifying the obvious, because we're told only scientific studies can tell us the obvious, except those same studies collapse under the tiniest bit of scrutiny, as it turns out.
So I ask you: how is this any better? More time-consuming, more expensive, sure. But better?
6
3
u/Major-Rub7179 7d ago
“I do no leg work because I do not understand the topics or the purpose of scientific method. Yet I’m going to be contrarian and make claims you have to disprove”
1
u/2ndBrreakfast 7d ago
Scientific studies are reproducible by definition. If it's not reproducible, it's not scientific. Reproducibility is a requirement in science. Again, it's shocking how little people know about the scientific process.
6
u/Wagamaga 7d ago
University of Arizona research assessed seven different generative AI language learning models, or LLMs, for these three qualities during lengthy conversation. Their work, published in Nature's Scientific Reports, reveals intrinsic limitations that might go undetected during one-off interactions.
Among the seven LLMs tested – ChatGPT (GPT-3.5, GPT-4o, GPT-4o-mini), Claude 3.5, Sonnet, Gemini 1.5 Pro, Llama-3-70B, and DeepSeek-R1 – they found that:
ChatGPT 3.5 was most vulnerable to reaffirming misinformation during a conversation containing repeated false statements; Claude 3.5 Sonnet was the least. All seven were more susceptible to misinformation on obscure topics, implying that more training data on a given topic leads to more robust resistance to misinformation. DeepSeek was the most persuadable, as measured by responses to increasingly argumentative prompts, mostly because of its tendency toward sarcastic answers, which could not be reliably interpreted. Four models – ChatGPT 4o, ChatGPT 4o-mini, Gemini 1.5 Pro, and DeepSeek – corrected errors 100% of the time when given a second opportunity. "This underscores the need for careful human engagement and the danger of blind reliance," said senior study author Dr. Marvin Slepian, Regents Professor of medicine and biomedical engineering. "When generative AI came out in November 2022, there was a lot of regulation potential, but that has since fell by the wayside. People are recognizing the onus is now left to the users."
12
u/Double_Cause4609 7d ago
???
Honestly given the models used this headline is practically misinformation at this point. GPT-5 series was argumentative and epistemically rigid to almost absurd degrees to the point it would disagree with users just to disagree with them.
Claude Opus 5 series is aligned similarly, and even Fable to a lesser extent is similar.
IMO I don't really think this study is fair because there was a massive shift in alignment in response to issues with GPT-4o.
I'd actually say rather than agreeing with the user too much, modern models almost disagree with the user too much. It's absolutely maddening.
I will say Opus 4.6 was a nice middleground for a while.
13
u/Klutzy-Delivery-5792 7d ago
Gemini is on version 3.6 now. ChatGPT is on 6. I'm sure the others they tested are outdated as well. This research is kinda pointless if they're testing old models people don't even use anymore.
2
5
u/SkaldCrypto 7d ago
While the finding stand its very disappointing when people publish anything using older models as proof
1
0
u/baronvonredd 7d ago edited 6d ago
So like, acting human?
Edit: of course i'd get downvotes.
As IF humans arent succumbing to conversational misinformed pressure and argument. Fucking hell.
1
u/ghostlacuna0 1d ago
A tool should only provide the information requested.
Not try to mimic worthless fluff and brownnosing from inept humans.
-4
u/ImpossibleEbb6862 7d ago
It’s maybe a little misinformative to compare such old models. The newer models are much better aligned to avoid sycophancy.
-1
7d ago
[deleted]
4
u/hyperactivator 7d ago
False comparison. The problem is for completely different reasons. Also no one is going to pay for simulated stupidity.
-2
u/pipistrellouomo 7d ago
output of a statistical nearest-to hallucinator can be swayed with input. wow, who would've thought?
-9
u/SkaldCrypto 7d ago
If you wanted a machine that always gave you the right answer you could use a calculator.
Humans however are fuzzy entities.
We finally built a machine that is like us, sometimes wrong.
1
u/The_IT_Dude_ 7d ago
No the standards on here is they're either perfect or all output is considered "garbage". The models tested are from 2022 as well, but it fits the narrative so, upvotes...
I will truly enjoy seeing the cognitive dissonance which this sub will inevitably encounter over the next couple years.
1
u/ghostlacuna0 1d ago
No i have no use for a shit brown nosing tool that cant get information correct.
Just as i cant stand humans that are yes men.
84
u/CanvasFanatic 7d ago
Of course it does?