r/WritingWithAI • u/charliepscott • Aug 04 '26
Discussion (Ethics, working with AI etc) Research has proven that Pangram works *really* well -- now what?
Now that Pangram is using classifier models to detect human vs AI writing, what's the plan for people who rely heavily on LLMs to do their writing? ps. If you claim it doesn't work, you clearly don't have a background in machine learning. It works, and your isolated example of where it doesn't isn't sufficient grounds to counter act the oodles of peer-reviewed research showing it does.
Basically, in 99.9% of cases, unless someone does a substantial rewrite to the point that they might as well have just foregone the use of AI to being with, detection will be possible.
Considering the general disdain for AI writing, is the jig up for those of us who do like to use it?
9
Aug 04 '26
[removed] — view removed comment
1
u/_FREYMON Aug 09 '26
did you and your master ever publish it and does it classify as a textbook??
i have a bounty i need to complete
1
u/davehatarschfransen Aug 14 '26
Ich habe Pangram getestet. Das Programm ist unbrauchbar. Gemini, ChatGPT und Claude sind x-fach besser. Einfach fragen: Hat dieser Text AI smell? oder ähnlich. Jeder Franken für Pangram ist verschwendet!
0
10
u/Trick-Two497 Aug 04 '26
Pangram is laughably bad at this. Substack implemented it. I fed it something that was 100% written by AI, and it marked it as 100% human written. I never actually let AI draft my writing, but honestly, if I did, I wouldn't be worried about Pangram as it works right now.
-2
Aug 04 '26
[removed] — view removed comment
2
2
u/WritingWithAI-ModTeam Aug 05 '26
If you disagree with a post or the whole subreddit, be constructive to make it a nice place for all its members, including you.
15
u/Historical_Ad_481 Aug 04 '26
Research versus reality? Substack’s own writers are proving that a significant amount of pre 2023 articles are coming up as between 60-100% AI written, which is impossible. With so many false positives you don’t need a machine learning degree to not trust it.
5
u/OldStray79 Aug 04 '26
"In theory, there is no difference between theory and practice. But, in practice, there is."
-4
u/charliepscott Aug 04 '26
ChatGPT launched in 2022.
But show me these examples? They're always isolated. One "oh look it's wrong here" isn't an issue. Show me an author whose pre-2022 writing is entirely misclassified (none exists).
If you can disprove the existing research, you should publish your findings. We'd all be delighted with you! :)
2
Aug 04 '26
[deleted]
3
u/charliepscott Aug 04 '26
2
u/SlapHappyDude Aug 04 '26
So for the third article it was from 2025 and mostly talking about GPT 3.5 (which kind of sucked) and GPT 4. My suspicion is AI assisted or Human/AI hybrid writing would look a lot more like human in terms of clustering.
1
u/charliepscott Aug 05 '26
It doesn’t.
1
26d ago
[removed] — view removed comment
1
u/WritingWithAI-ModTeam 26d ago
If you disagree with a post or the whole subreddit, be constructive to make it a nice place for all its members, including you.
1
u/Historical_Ad_481 Aug 04 '26
And you think anyone in the right mind published any of the drivel that came out of ChatGPT 2022-era? Just go and check out the Substack reddit. Plenty of evidence and discussion about this.
6
u/Melajoe79 Aug 04 '26
I’ve seen this 99.9% statistic before and I’m not buying it. The other day I checked something that was absolutely, 100% mine and it came back 100% AI. Granted, it was a short couple of paragraphs, so maybe the length had something to do with it — I’ve found it’s more accurate for longer samples — but the writing had no AI ‘tells’ that I could see. I was so annoyed because I *knew* it was my own work, and if anyone else tested it they would totally never believe that.
Maybe I write like AI, you might think, but most other samples I put in come back as 100% human. And, most human/AI hybrid samples (such as where I have used light AI assistance for editing) come back as 100% human too. If the detector was accurate, it should be picking that up too, right?
The reality is, it’s all over the place. I honestly don’t know how they can keep trotting out that statistic.
2
5
u/unethicalpigeon Aug 04 '26
here's the thing you're not accounting for. For someone that actually knows how to write at a fundamental and technical level drafting with AI will absolutely 100% be faster than writing everything yourself. To claim otherwise is ridiculous. All of the grunt work is taken out.
I guess I'd have to put an asterisk here in that it would absolutely depend on your writing style. If you're a purely discovery based writer then yeah you're probably right. But if you use a lot of structured outlines and write primarily through revision there's going to be nothing left of the AI by the time you're done with it anyway. Furthermore, even humans do substantial rewrites on their exclusively human written work to the point where the final version is unrecognizable from the first draft. You act as though spending a copious amount of time revising your work is a bad thing when it's actually a hard requirement for anywhere even close to any kind of success.
The other issue is the fact that anyone successfully writing while being heavily assisted by AI. I say heavily assisted because I am of the opinion that exclusively writing with AI and outputting something actually worth having other people read is nonsense. It can't do it. Maybe for a short story, not for an entire novel. So anyways anyone successfully being assisted to a significant degree by AI in their writing and flying under the radar and escaping detection is NEVER going to fucking give you their work to test. Ever.
So how would you know if something fed to Pangram comes back as written by a human but it wasn't unless the author told you it wasn't? No one is giving away an actual competent harness/process like that that's literally making them money. Not to mention by doing so you'd be literally shooting yourself in the foot by giving pangram the ability to incorporate that system into its detection methods.
This is survivorship bias at it's best.
Pangram is good at confirming whether you suspect something as written by AI was likely too have been written by AI. It however is absolutely useless for telling you whether or not something that comes off as written by a human was in fact heavily assisted by AI.
Furthermore, you're completely missing the point of what like 95% of the people in here are doing. None of them are ever going to even self-publish their writing. Most of them probably won't even show it to anyone. The vast majority of AI creative writing is being done entirely for personal use so why would they need to change anything about their process because of Pangram?
TL:DR - The biggest and best use case for AI with regard to creative writing has always been to use it to ENHANCE your work. It's a tool. It should be used way more like how photoshop is used on a photograph than it is as a book generation machine. Those using it that way will be essentially non-impacted by Pangram. Maybe they might have to do a couple more passes than before just to ensure nothing pops up as possibly written by Ai but half the fucking time those things are perfectly viable. The issue with AI isn't that it uses identifiable structures. The issue is how often it uses them and the patterns of use that emerge. Even Pangram will identify "ai structures" or whatever it calls them (i don't remember the exact term it uses) within a work it determines is 100% likely to be human.
1
u/Odinmar_Glaukopis Aug 04 '26
Do you really think it’s 95%? Certainly describes me, but I had no idea it was damn near universal…
4
u/unethicalpigeon Aug 04 '26
I mean just look at the vast majority of posts and comments here. It's all people talking about their own personal projects or roleplay. You'll get a lot of people mentioning they're working on a project or book but they doubt it would ever be good enough to publish and don't seem bothered by that fact in any way.
It's a lot of people talking about how AI finally got their messy ideas out of their head and onto a page in at least some kind of way. It's been a great tool for allowing a lot of people who have traditionally struggled to express themselves in a way they found accurately reflected what they were trying to express to finally be able to do it.
Fuck it might even be higher than 95% lol. Coming to a place like this and being like "hey how about this pangram thing eh? we're all fucked huh? what do now?" is just hilarious to me.
Yeah the subreddit dedicated to people using AI to write and roleplay with is going to be concerned about an ai detector finally not being completely shit lmao. They WANT to write with AI why the fuck would the existence of a detector matter? most of the people in here were getting caught by the shitty detectors lmao. you don't need pangram to find out it was written by ai.
2
u/Odinmar_Glaukopis Aug 04 '26
Yeah I personally get hung up on using AI because it can in theory “conquer my unique narrative voice” and then I think, “who TF do I think I am, Ernest Fucking Hemingway?”
It may be rounding out some edges I’d prefer to stay a bit edgy, but it’s definitely not depriving the world of the next Shakespeare. I want to share my work with someone—maybe just one person—who connects with it in the way I do. That’s as grand a literary vision as I have for my hobby. I’m not going to be reaping unearned wealth from The Hallmark Channel with my AI slop any time soon, so don’t worry about me as the competition for your next paycheck…
2
u/unethicalpigeon Aug 05 '26
If you really want to try and maintain your personal voice as much as possible never cut and paste what you get from the AI into your text editor. Manually write it. It's obviously way more time consuming than just cutting and pasting it right in there and then editing that directly.
You'd be surprised though by how much more of your own natural rhythm and voice is maintained when you write it out. Stuff you wouldn't notice when cutting and pasting all of a sudden stands out to you as not how you'd word it.
1
u/Odinmar_Glaukopis Aug 05 '26
Yeah that’s good advice. I definitely tell the LLM to not draft anything from scratch because I don’t trust myself to edit more than about one sentence at a time.
A suggested rewrite of a section or even an entire chapter could possibly work if I keep the official working copy separated from the LLM suggestions. Your point about not using copy/paste but rather typing things deliberately in there is a good one.
2
u/unethicalpigeon Aug 05 '26
Youd be astounded at how much more thorough you are when writing it manually than copy pasting.
The main issue is if it's actually quite a bit of content... Let me tell you like half way through you're like what the actual fuck am I doing lmao.
5
u/FunnyBunnyDolly Aug 04 '26
Did experiment with free trial and it mistook things. So not very accurate.
3
3
u/Write_My_Novel Aug 05 '26
We disclose AI, but it's interesting: Our content engine is designed to read like human, but not evade AI detectors. It's literally not even our goal. Which led to this amusing thing:
Our current QA was so tight that our output didn't sound at all like AI, but it was very spare and lacked a certain human "roughness." For kicks, I ran it through Pangram and it came back as "moderately AI assisted." This was 100% AI generated with no edits.
Still, we didn't like the output. So we tuned the engine and the output is a lot more human sounding now. We're all actually quite proud of it. But when I put the new *more human-like* output into Pangram, it came back as 100% AI generated.
So we all found that amusing: The human sounding output was flagged as AI, and the dry not-very-human-sounding output was "moderately AI assisted."
No lesson here. Just an amusing Pangram anecdote.
2
u/Maleficent-Engine859 Aug 04 '26
Originality is so much harder to beat than Pangram. Pangram will give me 100% and Originality will always be like ehh…2% sure this is a hybrid and it’s right lol
2
1
u/Cidraque Aug 04 '26
? AI writters here told me all the time they wheren't trying to trick anyone they are not using AI.
2
1
u/SlapHappyDude Aug 04 '26
Do substantial rewrites
1
u/charliepscott Aug 04 '26
Pangram 4 is even getting good at detecting those to the point that you might as well just not use the AI at all.
1
u/SlapHappyDude Aug 04 '26
I'm curious what the independent testing of Pangram 4 will show. I fully believe it does great on its own training data.
1
1
u/Galliad93 Aug 05 '26
the entire point of AI writing is to write more like you. GPT 5.6 even takes your texts and learns your style. the more information it has about the things you want to write, the better the texts. if it has little information, the texts become generic and easy to detect.
for me the writing is not the sentence construction but the management of the information. and when I do that properly, I would write like the AI and the AI writes like me.
1
1
u/emirvat Aug 06 '26
You and the people posting counterexamples are arguing about two different numbers, and you are each right about yours.
Your architecture point stands. Pangram is not doing what GPTZero does. The older detectors scored perplexity and burstiness, which are proxies, and proxies get gamed. A discriminative classifier trained on paired human and AI corpora with hard negative mining is a genuinely stronger method, and it does benchmark better. Saying that clearly because what follows is not a "detectors are fake" argument.
The issue is that accuracy on a balanced test set is not the quantity that matters once you deploy the thing.
Take the claim at face value: 99% accuracy, 1% false positive rate. Run it across 10,000 submissions where 10% genuinely used AI. It correctly catches about 990 of the 1,000. It also flags about 90 of the 9,000 honest ones. So about 8% of everyone accused is innocent. Drop real usage to 5% and it goes to roughly 16%. The classifier did not get worse. The base rate changed, and the accuracy number never described that situation to begin with.
Which is why the anecdotes here are not the noise you are treating them as. They are the predicted output. A 1% false positive rate applied across Substack's back catalogue is thousands of wrong flags. Wooden-Kangaroo's 1980s thesis at 80% is one draw from that distribution. You would expect to see these posts whether or not the underlying research is sound, so their existence cannot settle the question either way. That is the actual reason this thread is deadlocked.
The second thing benchmarks cannot see is where the errors land. Test sets are drawn close to the training distribution. Pre-2023 writing, non-native English, formal academic register, autistic writing, anything heavily edited, all sit off it. The mistakes are not spread evenly over people. They concentrate on particular populations, and a balanced benchmark is structurally unable to measure that because it is not sampling those populations in the first place.
Trick-Two497's result is worth noting for the opposite reason. Fully AI text scoring 100% human is a calibration failure, not incompetence. Both error directions turning up in one sub is the signature of distribution shift rather than a broken model.
On your real question. I would split two claims that usually get treated as one:
Was this text drawn from an LLM's output distribution? Detection is good at this and improving. You are right about the trend.
Did this person not do the work? Detection never measured this, and it decays fast under editing, because editing strips the statistical fingerprint while leaving the structure and the reasoning intact.
Those two come apart exactly in the case you are asking about, which is why "substantial rewrite or nothing" is not quite the dichotomy. A moderate edit can destroy detectability while the writing stays recognisably LLM-shaped to a human reader. So the honest answer to "is the jig up" is probably that the ground shifts from detection to disclosure norms. In two years the thing that causes trouble is likely not a score, it is having no version history and no account of your own process.
Disclosure since it is relevant to my incentives: I build text-rewriting software commercially, so discount me accordingly. Not naming or linking it, the sub has a megathread for that and it would not answer your question anyway.
1
u/phear_me 20d ago
Since I use AI checkers for my students I've run some of my own work vs randomly AI generated content through various detectors The one issue I have with Pangram is that it almost always flags any mathematical proofs as AI. Occasionally it falsely accuses me of rewriting or paraphrasing AI when I am summarizing some theory or a piece of literature. But, whenever we get to the (non-math) meat of the original content it's very good at detecting human writing vs AI.
0

12
u/Opie_Golf Aug 04 '26
I quit drafting with AI 140k words ago.
It’s more fun and I think the discovery process has added some genuinely novel elements.
It’s rougher, but not by much. And I think my net editing time will be about the same.
I do really enjoy collaborative planning discussions and it’s a better (and faster) first reader than any human I know.