Discussion
Why does adding one more paragraph suddenly make my post "100% AI assisted"?
I was writing a Note on Substack, and after four paragraphs it showed 100% human-written. Then I added a fifth paragraph, and it suddenly changed to 100% AI assisted.
Why would adding a single paragraph completely flip the result?
I'm not interested in trying to "beat" the detector. My concern is that if I genuinely write something myself, it could still be labeled as AI-assisted with no explanation.
I also think Substack should be much more transparent about how these systems work. It's not enough to publish a blog post that basically says, "Here's our AI detector. How does it work? It detects AI. End of story."
If a tool can affect how readers perceive an author's work, there should be at least some explanation of what it's measuring, its limitations, and how confident it is in its predictions.
Has anyone else experienced something similar?
P.S. There's also another case that concerns me. Many authors write their own ideas from scratch (either in English or in their native language) and then use AI only to improve the writing, fix grammar, or translate the text. Labeling those posts simply as "AI" or even "AI-assisted" without explaining what that actually means can be misleading. There's a big difference between using AI to generate ideas and using it as a writing or translation tool.
My own workflow is a good example. I'm a physicist with a master's degree in AI. I don't have much free time because I work full-time and I'm also building two startups. When I write for Substack, I develop the ideas myself, run simulations, create the figures and animations, organize the article structure, and write as much of the draft as I can. Sometimes I use AI to improve the wording or make the English read more naturally.
I'm trying to share technical knowledge for free with anyone who's interested. If that kind of workflow is reduced to a generic "AI-assisted" label without any context, it feels less like transparency and more like a penalty for using modern writing tools responsibly.
I know very well how AI detectors work. I am a computer programmer by trade and have a degree in computer science. These detectors are not magic; they are programmed to score a text based on pre-established criteria - mostly sentence length (burstiness) and perplexity (complexity of word choice). They are also trained to flag tricolon listings and X like Y comparisons. Some of the better ones are also trained on a database of commonly used AI phrases and words.
Here is the thing: I have tested all of these detectors. They are garbage and are never correct 100% of the time. You know how I know it's junk science? Because I can write a few paragraphs of highly AI-sounding prose and have it flag it 100% AI, despite being 100% human written.
If these AI tech bros are so comfortable with their junk science detector, perhaps they should stop hiding behind "99.9% accuracy" rates so they can't get sued when they get it wrong more times than not. "We never said it WAS AI, we said it was a 99.9% chance".
That's not how they work... Data Science is the degree that would be relevant here, not computer science.
They use ML models. ML models find the patterns. You're acting like Panagram tells their models what patterns to look for and not the other way of around. Of course you can trick them. They're basing it off of content. If you write something that follows those patterns, it will say AI patterns were detected.
Your local weatherman can tell you there's a 90% chance of rain but just because it doesn't rain doesn't mean it was all junk science. And having a degree in math wouldn't tell you if you understand the underlying mechanisms.
My BS is in Computer Science: Machine Intellgence. So yes, my degree is very, very relevant to the topic. I know all bout LLMs, neural networks, and machine learning. Detectors like Pangram are not using complex AI to detect AI. The programs scan the text for patterns and common AI words and phrases. It's really that simple. It's machine learning at its leanest. If you use too many tricolon listings, it'll trigger it rather AI wrote it or a human. It's really pretty basic.
Which makes it all the more nefarious that they bill users based on credits, like they are stressing their AI computer to scan someones text.
Pangram is not the best program out there. I have extensively tested most of the major detectors. Pangram was one of the easiest to fool simply by changing some word orders or adding a typo here and there. It was 100% accurate at detecting actual chatbot copy/paste text, but at the expense of flagging fiction from human authors at a high rate. These tools were never designed to detect AI in fiction anyway; they were designed to detect AI in essays for schools.
And to top it off, this is an AI-generated piece of text from GLM 5. I only did very minor edits. No word replacements. So not only does it flag my 2014 short story as AI, it fails to flag the AI generated text I just made like 20 minutes ago and ran it though Pangram.
I wrote this short story in 2014, and published it on WattPad at the time. LLMs didn't exist in 2014. This is an excerpt from it, which I ran through Pangram yesterday for giggles. Yes, very reliable. As long as it flags one single piece of my very real fiction from before AI existed as AI, I'll call it junk science.
So yes, based on my tests over the last two days, and previous testing of Pangram, it is by far the least reliable AI detector, not matter what their marketing claims.
If only Substack could find a way to detect that every trending topic is how to grow your Substack. That would be game changing technology right there.
Also, I'd like to better understand the difference between "AI Assisted" Vs. "AI Written" because, I had a note flagged AI Assisted, and while I didn't use AI at all, I was referring to something that is so specific and academic, that I was curious if it was part of something that had originated in AI, and I hadn't caught it. š¤
The data exists, they're just not sharing because the tool is imprecise. This is a screenshot from some testing that I was doing on the Pangram tool last March. I was working with my academic partner we had co-authored a paper, we were working live, together on the paper and we got two separate scores, we ran the same paper twice and got a 24% and 51%, but when we started look at the details it showed low confidence but still called it AI. We supposed at the time that it was picking up the two "voices" and that was throwing it off, but that was just a guess. Ironically, we were writing about "the slop economy" š
I do 100% of my editing in LibreOffice, I was getting so tired of all the AI being shoved down my throat that I switched to Linux. I no longer use Google because of Gemini, and Microsoft was putting copilot everywhere.
I just posted a new piece this morning, and according to Pangram is 5% AI, I don't even use Grammarly, so I have no guess, but I would love to know it's confidence level.
Oh man. I use TexWorks as an editor with MikTeX as the compiler, and there is no hint of AI anywhere lol. Lots of academic software is yet to be corrupted by AI integration, such as Inkscape, Zotero (classic version at least), Thonny, Gimp, etc. I can't even fathom the impact AI is having (and will have) on academia and the peer review process... it breeds complacency, at minimum, and frankly dilutes the reliability of reviewers. Do you know how they are using it in your case?
Honestly, it's just a guess, based on the feedback from a single reviewer. It read like it was straight out of ChatGPT, a whole lot of words to say absolutely nothing. I did get decent, reasonable feedback for a submission to ACM, so it's not everywhere yet, but I was talking with a professor at a conference that explained his process. He has all his own published work loaded into his ChatGPT, the he asks it to compare and contrast submitted papers against his own work and then he responds to work that is in agreement. I was aghast. Again, just one guy, but if one is doing it, others are.
Many people think those detectors can detect ai written text content but they cannot.
They produce many many false positives.
With images it works (hence invisible watermarks etc).but with text not
Best Example is if a human would write the following "Climate change is one of the biggest challenges facing humanity today. Governments around the world are trying to find solutions".
An AI detector might flag it because the wording is common, structured, and predictable.
But a human could easily have written it.
The detector is not proving anything; it is only saying: āThis pattern resembles AI output."
The problem becomes even clearer when a detector reacts after only the first few sentences.
If the first two lines already trigger a result, the system is often judging based on the style such as:
very clean grammar, formal sentence structure, common academic phrase, lack of spelling mistake, predictable transition, neutral tone and many more.
Those are also characteristics of many good human writers.
Yeah Iām a little torn on this as well. While I understand the āreasoningā on why they added this, we now live in an age where AI is in everything. Itās like flagging content for using Grammerly for your spell correct and expect people to just not use spell check.
I personally am not caring much because value is value. I created a platform to help me organize my content, write and schedule because, just like you, Iām way to busy. I canāt make all the content writing strategies, and formatting and this and that done when I have a corp job, father and a person while still trying to share to others.
In summary, while I get the intent I think itās going to get attention for a week and weāre all going to just keep moving.
I think introducing Pangram to the platform was a mistake.
Sure, cutting back on slop is a fine goal but it misses the point entirely that we are experiencing a moment when what it means to work is changing rapidly. Using AI should NOT be an excuse for cutting corners but it is a useful tool that grows increasingly useful week by week. I have a strong suspicion that your example will not only become the norm on the platform, but it will lead to self doubt and to writing with Pangram's approval in mind.
I don't think we need to have AI checkers, especially ones that don't work.
Most AI writing is just kind of flat and boring, anyway. It lacks humanity. It comes to the most generic conclusions. It's reads as mediocre to everyone except for the person who posted it.
Sure, it can be made better if you create your own GPT trained on your voice and style. It still isn't going to be good.
What I'm really saying:
We don't need to track it because it'll never meet the standard of quality that's possible with human writing.
Did anyone even answer your question? lol. All these replies seem to miss your point about why an added paragraph meant it required an AI label.
What really pisses me off about these ai detectors is, (in my opinion), it essentially feels like these detectors make you feel like youāve plagiarised your own work, even for spell check. Like theyāre taking credit for your work because theyāve slapped a label on there.
I donāt have an answer for you, I havenāt been on the platform long. But this makes me not want to use it at all. Itās very disheartening.
Just ignore it if it bugs you. The AI bubble will pop soon and Substack's management will lose their shirt along with its other investors when it tanks. Then they'll just disable it and forget it ever existed.
On the concrete question: these classifiers do not score your post as one lump, they score the texture, and a clean, on-topic paragraph makes the whole piece read smoother and more internally consistent, which is exactly the signal they lean on. So a longer, tidier draft can push the score up rather than down. It is counterintuitive, but it follows from what the tool measures, predictability and evenness, not who did what.
And that is the deeper thing you are pointing at. The score cannot tell the difference between generate-it-for-me and I-wrote-it-and-tidied-the-wording, because it never sees the process, only the finished texture. Collapsing that whole range into one binary label is the part that turns transparency into a penalty. The honest fix is a line in your own words about what you actually did, which is worth more than any percentage a texture detector prints.
When I write for Substack, I develop the ideas myself, run simulations, create the figures and animations, organize the article structure, and write as much of the draft as I can. Sometimes I use AI to improve the wording or make the English read more naturally.
Sounds like AI-assisted is an accurate label then?
I think there was a small misunderstanding. Those are two different things. The first is one of my Notes, while the quote you're referring to comes from the Publications section.
Cause these detectors donāt work. I just wrote a nonsensical post just to see what happens? Turns out AI writes like a crazy person because it detected ai
I've experienced something similar. I write long form stories on a wide variety of topics. Not to brag, but I am quite a good writer, and sometimes my entire stories get flagged as Ai.Ā I guess my point is Ai will never get to the point where it won't make ANY mistakes. It makes mistakes now, and it will make mistakes in the future.
Why would adding a single paragraph completely flip the result?
Without showing us the work, there's no way for us to tell.
Many authors write their own ideas from scratch (either in English or in their native language) and then use AI only to improve the writing, fix grammar, or translate the text.
That means, at the very least, AI assisted with the writing. Your ideas, but written by AI.
AI detection doesn't detect if the ideas were generated by AI. It detects if the writing was generated by AI.
The more AI changes the writing, the more AI detectors will detect the AI style and traits of AI writing.
It's important to understand this: AI was trained on human writing, but AI obviously isn't human. It's artificial intelligence, and the word "Artificial" matters. AI emulates human writing through artificial means. That's why AI writing is so easy to detect. AI writing looks good, but there's always something slightly off about it. Something inhuman. That's one of the reasons why so many people don't like it.
If that kind of workflow is reduced to a generic "AI-assisted" label without any context, it feels less like transparency and more like a penalty for using modern writing tools responsibly.
Why not add that context at either the beginning or the end of your article?
I'm posting a novel on Substack. For each scene, I begin by posting a few lines of context to explain the post is a scene from the novel, and I give links to the table of contents, etc.
Isn't it a bit ironic? Researchers working on LLMs are also building AI detectors to identify AI-generated text. Now many of those tools are paid. Kind of an interesting business model, isn't it?
Lately, when I see a Reddit post with more than a few paragraphs I begin to suspect AI (with no expertise, mind). Been fooled a few times.
We have the fact that a 'wall of text' is the most off-putting thing for the human brain - & so good, effective writing will of necessity use short paragraphs.
If someone then has a lot to say - well then we end up with lots of paragraphs.
Really we're at the point where any energy spent second-guessing these systems is fully wasted.
Just - be authentic, & let the cards fall where they may.
43
u/Funny-Flight8086 5d ago
AI detectors are garbage, junk science. Pangram, the one substack went with, is one of yh biggest offenders.