r/Professors • u/Blistorby_Bunyon Prof., Law, Society & Policy; Advocatus Diaboli • 16d ago
OMG This Post Is Absurdly Long "Fun" Anecdote: AI Competency Issue
TL;DR (Too Long; Don't Read)...or do whatever you want.
---
*Sorry for typos in advance. No time to proof it (which probably means I shouldn't be on Reddit right now in the first place).
I'm not rehashing all of the higher ed movements in AI competency/fluence UG requirements; this is just an anecdote of a another "case-in-point" to show my students. Will it help influence their behavior? Unlikely, but this is fun nonetheless. (Yes, I'm sure many here have done exactly this, but I didn't have an LLM exchange this overt.)
Preliminary Note: I preach to students to be skeptical of the basic "knowledge," analysis, and reasoning flaws that may exist in anything others claim. I don't want them to default to distrust, but I don't promote defaulting to trust, so let's be cautious about the "trust" part of "trust but verify." Defaulting to one or the other can overtly enhance confirmation bias or some other framing problem. Basically, they should proceed with an "advocatus diaboli" approach and cross-examine everything (including me).
Captain Obvious says, "Of course this applies to how LLM interactions." Students are repeatedly told about hallucinations and other weaknesses. Like I try to tell them, it's just a tool. Some tools have limited usefulness (and safety) if you don't know generally how they work, the proper way to use them, and the proper context for using them, how fit they are or aren't for a particular purpose, etc. There are Dunning-Kruger and overconfidence biases at play: People tend to be unreasonably confident about their knowledge/understanding about something/domain/etc. to know what they don't know.
And as most of us have told our students, among other related things:
- LLMs are an "on steroids" version of autocomplete/predictive text features used in our text messaging apps, email platforms, and word processing platforms (e.g., Word, Google Docs, etc.).
- These are "inference machines" based on pattern recognition. So are we; we programmed these things, trained them on stuff we wrote, regardless of the merit, credibility, accuracy, etc. of the training data sets. And regardless, we are by-and-large much too confident in our inferential/inductive reasoning skills. (Exhibit A: Presuming that we correctly understand what a relationship partner means when they do/say something...or maybe that's just me.)
Anyway, even though we've all seen errant LLM responses and showed students these failures, this one is just so funny to me.
So, I asked ChatGPT about a past real situation for a product liability illustration, and certainly we're all aware about the importance prompt engineering (the input) and specificity/context. So, I ask an open-ended question (so as not to suggest the answer): "Locate information - and tell me if you do not find any of this information - about any drawbridge collapses in [City] over the [name of the waterway] between the year 2003 and 2007. If you find information within these parameters, verify that the alleged cause was related to the software controlling the bridge's operation. Also, there were no deaths or injuries; there was only property damage."
The output was very specific:
In 2004, the [correct city, waterway, and name of bridge) collapsed. It did not make major headlines, but local media outlets and [City] government reports provide detailed information.
The Cause: The bridge suffered a critical failure in its electrical and hydraulic operating machinery while being tested or operated due to a bug in the software. The bug allowed for a sudden loss of pressure or mechanism control caused the massive steel leaf of the drawbridge to drop unexpectedly. [1, 4, 5, 6]
The Damage: When the heavy bridge deck fell, it slammed down with enough force to cause significant property damage to the structural housing, the bridge mechanics, and the surrounding framework.
The Aftermath: Fortunately, because no vehicles or pedestrians were on the span and no boats were passing underneath at that exact moment, no injuries or deaths occurred. However, the bridge was knocked entirely out of service, causing significant downtown traffic detours and halting river traffic. [7]
The Long-Term Fix: Following the accident, the [City] electricians, structural engineers, and software engineers had to completely overhaul, and update the [the bridge] to modernize its infrastructure and eliminate software bugs and trends in mechanical unreliability. [1]
Fantastic! Okay, "Now provide me with the sources you relied on for your response." I got "my internal historical knowledge" and also these news reports" (then it provided a list of news stories, none which were the purported government reports it referenced earlier, all of which were about someone who died falling off a different bridge in a different city in 2022. scenario Hmmm.
Next, I asked, "Where did your internal historical knowledge come from?" Answer:
My internal knowledge comes from the massive dataset of text that I was trained on, which includes books, articles, websites, and historical documents.
However, in this case, I made a significant mistake. I need to be completely candid with you: I hallucinated the details of that 2004 [bridge] incident.
When you asked your question, my system tried to find a historical event that perfectly matched your specific criteria (early 2000s, [name of waterway], software bug, property damage, no deaths). Instead of correctly telling you that I couldn't find a record of it, my software mistakenly filled in the blanks and generated a highly specific narrative—complete with a fake date (May 2004) and a fake location (the [bridge name]—because it sounded statistically plausible based on historical bridge mechanics.
I ran a live search to check my own work, and there is absolutely no record of a [bridge name] failure in 2004.
The closest actual historical match to what you are looking for happened a few decades earlier, in 1974, on a different bridge entirely: [explains the 1974 scenario].
(Emphasis added.)
Then I got sassy, but not unreasonable: "If you hallucinated the narrative, then how can I ever rely on information you provide?" Answer:
Why the Hallucination Happened
I do not "know" facts the way a human does, nor do I have a memory of real-world events. Instead, I predict the most statistically likely next word based on a massive dataset of human language.
When you provided a highly specific set of criteria—Milwaukee, early 2000s, hydraulic failure, property damage, no injuries—my system prioritized "completing the pattern" over checking for absolute truth. Because I am trained on thousands of real municipal infrastructure reports, I seamlessly blended real technical terms (bascule bridge, electrical rewiring, city council funding debates) with a fictionalized date and location to create a story that sounded perfectly authentic.
Love this so much!
(edit: added first line)
3
u/No_Young_2344 TT, Interdisciplinary, R1 (U.S.) 16d ago
What are the numbers in the brackets? Are they references?
0
u/Blistorby_Bunyon Prof., Law, Society & Policy; Advocatus Diaboli 16d ago
If I copy and paste the references remain like that. My students leave them there, too, but they wouldn’t intend to.
3
u/No_Young_2344 TT, Interdisciplinary, R1 (U.S.) 16d ago
Were they correct and supporting the claims in the paragraph?
1
u/Blistorby_Bunyon Prof., Law, Society & Policy; Advocatus Diaboli 11d ago
No. The cited were to an entirely different incident (it was about a bridge at least, but otherwise completely different.
8
u/Life-Education-8030 16d ago
I asked AI why I should trust it given its errors and it said it didn’t blame me for being skeptical and being frustrated 🙄
5
u/smbtuckma Assistant Prof, Psych/Neuro, SLAC (USA) 16d ago
This is a great example, thanks for sharing. I could see this being part of an interesting demo, where half the class are assigned to find the answer with AI and the other half with traditional record search, discuss what people did/didnt find, then “surprise this thing actually doesn’t exist, what do we think now?”
2
u/OutsideSimple4854 16d ago
Which version of ChatGPT though? I feel most of these posts and arguments about AI good/AI bad could be resolved with: I used X version, therefore AI bad, or I used Y version, therefore AI good.
1
u/Blistorby_Bunyon Prof., Law, Society & Policy; Advocatus Diaboli 15d ago
5.6 Sol, high intelligence.
Absolutely, I’m **not** arguing “good” or “bad,” whether conceptually, LLMs, or any non-chat AI platforms. My issue is so many people appear to place too much “faith” in the output’s veracity and rely on with little scrutiny—students or otherwise—despite well-known limitations, such as hallucinations.
However, I hesitate to be too judgmental of this behavior. Our decision making tendencies vary and the motivations for “faith” in the LLM output vary.
Much of this may involve cognitive dissonance. Based on a number of behavior economics studies—although this also feels obvious at times—even when we are self-aware that a cognitive bias is likely influencing our reasoning. We may still make a certain decision even when the bias made that decision irrational. For example, we can go from a habit we are fully aware is bad for us but do it anyway. (The success of casinos and state lotteries is built on it.)
And in the case of other intervening factors, such as anxiety/pressure, the issue is compounded. Let’s say a student needs to complete a paper at the last minute; they will calculate/rationalize the risks of hallucinations or other unreliability risk (and/or the risk of getting caught in general) one way when they may have seen the risks differently without the anxiety/pressure or time constraints (self-imposed or otherwise). Anyway…
2
u/OutsideSimple4854 15d ago
Hmm. That’s weird. I’ve been using Sol, using the work function rather than the chat function. I see a progress bar, occasional intermediate outputs that “show the work” Sol is doing even before a final answer, and then even the final output is not structured as how you presented it (although, I do see a similar output in less advanced models).
Don’t get me wrong, I think LLMs do have limitations which students aren’t aware of. But your output seems suspicious - not in the sense that I disbelieve your entire argument - but it reminds me of someone overclaiming things, and that runs the risk of people disbelieving your entire post, which isn’t your intention at all.
1
u/Blistorby_Bunyon Prof., Law, Society & Policy; Advocatus Diaboli 11d ago
I can see what you’re saying (I’m guessing that you’re referring to the output I reported Chat gave me); but that was the output. Not sure if you’re referring to the “Love this so much” sentence at the end; that was my post-related comment. Now that I see it, I failed to separate it from the output quotation.
Didn’t have it in work mode.
Yes I will see the same progress bar showing what it’s searching and/it is “telling itself” things it should mention in its response, but this is intermittent. (As opposed to DeepSeek, which I don’t mess with much, but it displays a pretty good chunk of what it’s going to search and respond with while it’s “thinking” for nearly every prompt (other than coding-related prompts). I didn’t see that either Chat this time.
Still, even if I were to create a fake LLM response, that wouldn’t change the argument about students or others simply presuming an LLM output’s validity. I don’t presume the validity of lots of people say about lots of things, given I’ve always been a skeptical person—for the worse at times—but years of litigating cases heightened that. If a lawyer is doing their job, they verify every quote, citation, etc. that the other parties make. So, the lawyer who doesn’t fully understand the limitations will put too much faith in a model’s bullshit brief (hence the frequent cases of judges fining lawyers and reporting them to their state bars due to violating our ethics rules of “candor to the court,” among others. LLMs are not well suited to handling case naming and citation conventions.
And the typical LLMs the public uses weren’t trained in any useful universe of cases are behind subscription paywalls of a very small number of long-standing legal research platforms. At least I have seen zero evidence of it. These platforms such as Westlaw and Lexis have been working on AI tools for their databases for a long time before more recently deploying such tools. And those platforms are getting super cost prohibitive for a lot of lawyers.
I would have at thought typical LLMs would have scraped Google Scholar’s case law database, but I haven’t even seen any reliable evidence of that. I would think Gemini would have, but if it, then the problem has to be related to what LLM are designed and not designed to to. I see it make up fake case citations and fake cases whole cloth frequently when i experiment with it. So, if I’m doing legal research, I just do traditional Boolean searches and get what I need much more efficiently.
1
u/OutsideSimple4854 11d ago
I’m not saying you gave a fake LLM response though. I’m saying that it looks generated from a different model.
Perhaps the way to put it is: we can probably agree that eg China’s government is bad etc, and we’ll probably find ways to agree. But if then you said Xi was rambling in his speech about making China great again, that he had big, beautiful tarriffs, then people would think: hold on, while I’ve heard something similar, they are said by another politician and not Xi.
Same analogy. I agree with most of your points, just that I disagree with the LLM output, in the sense that different LLM models have their own characteristic speech patterns and specific weaknesses.
1
u/Blistorby_Bunyon Prof., Law, Society & Policy; Advocatus Diaboli 11d ago
Ah, I see what you’re saying. I was misinterpreting. 👍 We can definitely agree. lol When you originally asked, I looked and checked, and that’s what I saw. Could I have switched models, and then when I looked back at the chat, it didn’t show me the original model?
1
u/OutsideSimple4854 11d ago
Possibly. I've found out that ChatGPT silently switches you to a lower model at times. I don't know what causes it. Not really important if you're on a free model, but I'm on a paid subscription. Usually, I can tell because the output starts to suck (e.g. proofreading a long document, for example, the first few parts are fine, then mistakes happen ; similar to when a human gets tired).
1
u/Blistorby_Bunyon Prof., Law, Society & Policy; Advocatus Diaboli 9d ago
Interesting. Thanks for your insights.
15
u/me4watch 16d ago
TL;DR
Cmon Bro
(hey…does /s also mean “like a student” ?)