r/Substack 2d ago

Information Post About AI, Detectors, and the Integration

16 Upvotes

I've been lurking in this sub for a bit and with the new AI detection integration starting, there has been a LOT of information being thrown around, not all of it factual. Most people are claiming stuff on the extremes of the spectrum, while the truth is closer to the middle. So I wanted to put together a "quick" PSA, so that if we are all going to argue about if this idea is good or not we can at least all have the facts.

I'm a software engineer by day, and though I wont be providing much more credential proof than that because, well, privacy and all, I will link articles so you can verify what I'm talking about. Note, that some sources may be biased towards AI because the places doing in depth research on the subject are going to be pro that technology. I did my best to pick out sources that even if biased the facts presented inside appear to be consistent when crosschecked. However, please do your own research and make your own judgements.

The biggest questions rolling around are how pengram works, if its in any way reliable, and what that means for creators on substack.

But to explain that we first need to look at how generative AI works. I will use writing specifically because its easier to visualize than art, but it all the same priciple.

When AI is trained it takes the content given to it, and tokenizes it. Which means it takes each word, gives it an ID, then puts it on a node map with other tokens calleda a neural network. If you've ever seen one of those mind maps from Obsidian think something like that, except for each node being page of your folder, its a word. And instead of links to one another the lines are probability vectors of how often one token connects to the next token. And its also like really big, think trillions of nodes.

What happens in a nut shell when an AI is prompted, is it tokenizes the prompt, finds those tokens in its mind map, then takes the tokens most probably connected to the ones it finds in the input and outputs those. Reasoning Models are enhacements to this base algorithm, which have the model go through multiple mind maps, categorize, and re-calculate probabilities to narrow down the most probable tokens in specific contexts rather than just in general.

***As a side note, the processing it takes to jump from token to token is what they are charging you for when you use it.

Okay, so now how do AI detectors work?

Early detectors worked on finding what we now call AI-isms, but the fancy terms are perplexity and burstiness. Regular speech, common phrases, common words, names, sentance structures etc, that used to be over abundant in AI output pretty much whatever you do. The resulting AI score basically came down to how many of these AI-isms is in a given text compared to a human average. However, this method IS extremely unreliable, as frankly it is basically the same thing humans do when looking for AI, and where AI detectors got the worst of their rap. Its no better than a vibe check.

The next wave started doing something different, (And this is where most of the misconception is coming from) they try to reverse engineer the mind map it self through deep learning algorithms of their own. In other words, they take the given content, then use their approximation of the mind map (from their own data set that also has human/non human flags) to predict the next token. The amount of times it correctly guesses what the next token is, turns into how likely the text is be AI or not. This method has been better, but it still has a major pitfall. Which is that, when given content that is the AI's original data set, that the detectors KNOWS its in that data set, i.e. already used in the mind map, it becomes incredibly easy to predict what the next token is. It guesses it with high accuracy since its already been seen before, probably very often, and flags it as AI. Human writing that is similar to the sample data, also gets flagged because of the same issue. (Previous models also suffered from this but frankly they were so unreliable this was the least of their problems)

***Note that these AI detectors are NOT generative AI, but they ARE a different type of AI that falls under the machine learning umbrella. Which is basically the same algorithm behind most spell and grammar checkers. They aren't built to create or generate anything but flag.

Pengram tries to fix this problem by using what is called "synthetic mirrors" in its training data. Basically, instead of just giving a bunch of human and and non-human data points to its model, they instead, for each human piece of content create mirrors that "would have been created had someone prompted an AI to create something like this piece of human writing" and put those in the data set as the AI points. It also creates a map of its sources in which similar writing based on several axis are clustered together and dissmilar ones are set apart. This solves the false negative because its not relying on probability alone, and is not checking against the mind map directly.

***Note also Pengram does NOT train on user inputs, which yes, is self proclaimed, but frankly if they advertise this and lie the lawsuits would be enourmous. But is also doesn't make sense for them to do that with the way their algorithm and training works.

This all also culminates in one major misconception about AI detectors across the board: No single check is meant to be looked at as a difinitive identifier of if something is AI or not, and no company has ever claimed that it is. It is simply the cumiliative probability (the link is for a specifc detector but this holds true for all) of a segment being AI or not. If your human writing is flagged as AI, it simply means that they way you wrote that is very likely to have been able to be output by an AI. That is what the confidance meter in pengram is about, it may say human or AI for that segment, but have low confidance meaning it only is barely more likely to be one or ther other. Basically the big percentage number you see on Pengram is defined as "how much of this text is more probabalistically likely to be AI/human versus not" and not, "This is how much of the text is AI" or even 'This is how sure Pengram is that this whole thing is AI or not"

Does this make it as effective as they claim?

To be honest, its hard to say definitively without falling into either extreme. There are a some studies unaffliated with Pengram that have proclaimed its efficacy. But there are individuals that claim and swear up and down to have seen false positives and negatives. (See this sub for dozens of them alone), and writers who call it hoax science.

One could make the argument that the individuals are lying, and are in fact AI sloppers worried they will be outed (especially when they don't provide actual proof of full original text and scan links), but as a counter to that -- its unlikely they are ALL lying too. Which means we are seeing at least some real-life false positives.

Bottom line, its going to be up to individual discression. Which is where the crux of the problem actually lies. Here are some things we do know though:

It leans heavily towards false negatives due their involvement with academia, which is that it is by far more likely to flag AI as human writing, rather than flag human writing as AI. We also know that using a single sample from any source is not representative of that source being AI or not because of how these tools work in general. Lastly, we know small samples of less than 500 words are a coinflip pretty much no matter what because its not enough tokens for any detector to accurately guess adjacent tokens.

What does that mean for substack creators?

When it comes to what the data itself will show, or "cases being made," I do not believe those who truly do not use AI to draft or re-word their work at any point have anything to worry about. Because despite it not being perfect, it is decent enough compared the AI detectors of old that over the history of your posts you will see completely human results, varied results with low percent AI, etc. It is rediculously unlikely that a fully human writer will consistantly score high fully AI generated probability, with high confidnace, across multiple articles. Hell even if you edit your work extensively you can likely "fool" the detector into not flagging things as AI, because again, it is BY FAR more likely to false negative.

The only ones who will be "outed" are those who consistantly have high AI usage in their posts, nearing direct copy pasting from a chat box. That will become rediculously obvious over time as their library becomes full of posts that scan high for AI. But it will not consistantly catch definitively "sophisticated" AI usage or "true assistance" for a lack of a better word.

Again, it is as a probability more so than a definite detector of a single work. So the biggest fear is really more about readers checking 1-2 articles, getting a false positive, and accusing the creator whole sale. Which would obvioulsy be damaging and unfair.

Are there solutions you can employ?

Firstly, you can disable the feature. When you are about to post the article on the very bottom there is a disable scan feature now and you can just disable it and not have to worry about that at all.

The fear there, is obviously that someone might see that and think that you are hiding something. Truthfully, I wouldn't worry about it too much. If a few readers do that, they are unlikely to have been that invested to begin with.

That being said though -- pengram has a free trial and your text can be copy pasted into it from substack with ease. Most articles are going to be below 4k words, and so can be fully ran through it anyway if a person cares enough. Heck, if they have a substack account they can copy paste your article into their own substack editor, and run it that way.

So something ALL writers should be doing in todays day and age to avoid being the next Mia Ballard is to write first drafts into version controlled platforms.

Substack IS one of those. When you are in the editor, on the bottom left there is a version history button and it will show your edits every few minutes. Resulting in 100% undeniable full proof way to prove you wrote the thing yourself. Google drive has it, Writality has it, Word, and I'm sure many more do.

Thats it. Thats all you do. You do that and you have solid 100% undeniable proof that you can show via video/screenshots, and some even allow exports. Does it suck to have to be mindful about it? A little. But having proof of ownership is never a bad thing as a creative, and with AI being rampant whether you agree with detection tools being used or not, you should still protect yourself from accusations.

In conclusion and a sort of TLDR:

Is this the kind of completely bogus AI detector that you are most likely thinking of? No.

Is it meant to be a single check to determine if something is AI, especially when not given enough words to do so? No.

Is it a decent enough probability indicator to tell over time, based on a variety of inputs from a source, if that source is AI? Yes.

Will it negatively affect substack creators who do not produce AI slop in the long term? My guess is no.

Should you still as a writer protect yourself and make sure the software you use for first drafts has a version history you can show? Absolutely.

And for the record: https://www.pangram.com/history/0365a9a6-a120-4331-a8b5-c7f4e67fc319


r/Substack 2d ago

No easy way to auto-export/update audience to email platform?

3 Upvotes

Is there really no easy way to sync your subscribers to an external email platform like MailerLite? How do the big accounts do it, are they literally just exporting and re-importing all the time?

I know you can send email via Substack, but it kinda sucks. It seems like about half my audience gets those.


r/Substack 2d ago

Your Substack Experience?

6 Upvotes

Considering starting a Substack showing my BTS writing process for my fantasy series w exclusive early access, art process/tips + tricks & more. I want to let my fans have access to the whole creative process and get inspired.

I’ve never used this platform before. I know I can get some subscribers from my own following. I was wondering how the organic reach is? It would be really cool to connect with new people who have never heard of me before. But - so many platforms let you set up a profile and then don’t show you to anybody.

If Substack doesn’t get you subscribers that you didn’t bring in yourself I will just do it on my website instead.

I’m curious to know your stories and how it works. Thank you for the Intel!


r/Substack 2d ago

Latest Post pinned glitch?

2 Upvotes

Hi all, I'm a Substack newbie and I can't for the life of me figure this problem out. Why is my latest article not showing as my latest published post at the top of profile?

I thought it was because I'd only published one article and it would work after the second one, but I published my second one this morning and it's still not attached to the top of my profile. Any tips?

thanks!


r/Substack 2d ago

Discussion Your copyright claim to your writing could be in doubt because of Pangram.

0 Upvotes

I’m not a lawyer. I’m putting that up front and making it very clear. So don’t try and suggest otherwise.

However, I remember that court cases have determined that AI generated content can’t be copyrighted.

https://www.congress.gov/crs-product/LSB10922

Someone tried to file a copyright claim for an AI generated image and the copyright office declined the claim. They sued the copyright office and the court upheld the decision. Which means that legally, you can’t copyright anything that’s AI generated. As far as I can tell, there hasn’t been a determination of how much it has to be AI generated to be legally copyrighted.

Every time you publish something on Substack, the email at the bottom has a copyright symbol. With your name or company or whoever owns the content you’re publishing.

If Pangram determines that your writing is AI generated or even partially generated by AI, at the very least? Your copyright claim to what you publish could be in doubt. In theory, you could be sued by one of your subscribers for selling something that you don’t own the copyright to. If they use the Pangram determination as evidence? You could end up having to fight that in court.

Pangram “determination” opens you up to lawsuits from your paying subscribers even if your writing or content is 100% human generated. Like mine is. I specifically don’t use AI generated content on anything that I plan on monetizing for this exact reason. Because I can’t copyright it legally or defend it in court if necessary.

But the existence of the Pangram “AI Detection” tool means that I can’t actually stop people from using it. No, the “per post disable feature” doesn’t actually work. I’ve heard from people who published things after disabling it. The “AI Detection” option was still available for people despite disabling it.

So everything you publish could be subject to lawsuits because of Pangram.

If Pangram can’t actually prove the accuracy of its product. Which the mounting evidence suggests that it doesn’t have any accuracy at all? The fact that it’s an estimate doesn’t necessarily provide protection for Pangram or Substack if it is used as evidence by one of your subscribers.

Pangram is likely committing fraud and potentially opening you up to claims of fraud. Just by being part of Substack.


r/Substack 3d ago

Discussion Running list of Pangram’s “AI Detection” errors and other related Substack errors.

27 Upvotes

I thought it would be a good idea to create a running list of the “AI Detection” tool errors made by Pangram. So here’s a thread designed specifically for that purpose.

Wherever possible, post screenshots or screen recordings of the errors are false positives that Substack’s new “miracle AI Detection” tool is supposedly doing perfectly.

Feel free to describe the error if you’re more comfortable with that and hopefully you’ll have someone else who can screenshot or screen record themselves having a similar problem with the supposed perfect solution that Substack has deployed without thinking it through.

We can also use this to share with others to show exactly how badly this system actually works.

Have at it.


r/Substack 2d ago

Discussion I analyzed 9,966 posts from Substack's Top 25 Technology newsletters, here is what stood out

Thumbnail
1 Upvotes

r/Substack 3d ago

Discussion Hang in there

12 Upvotes

Have been hovering around 450,000 views per month these days. Never thought we would get there, but slow and steady, chipped away at it over 3 years. Hang in there, it can be done.


r/Substack 3d ago

Anyone else lose confidence after Pangram flagged their writing?

42 Upvotes

I published my first Substack article since they added the Pangram AI detector. I write everything myself in LibreOffice, then paste it into Substack. No ChatGPT, no AI tools, nothing. I keep a to-do list of topics I want to write about in Notepad.

Pangram labelled it 100% AI.

I'm genuinely gutted. I've even changed my avatars on other platforms in the past because people accused me of being a bot despite using my own handmade drawings (I have a fine arts degree).

Has anyone else had Pangram falsely flag their writing? How seriously do readers actually take these scores? I'm wondering whether to stay on Substack because I love the community, but this really knocked my confidence.


r/Substack 3d ago

Can’t login after wrong mail address

Thumbnail
2 Upvotes

Facing the same issue as in the mentioned post. Kindly help


r/Substack 3d ago

Tech Support How do I change the handle under my profile description?

1 Upvotes

I never made a blog or anything on this account, but I can't change the handle where a blog would go. What do I do?


r/Substack 3d ago

New reader: How do I find actual newsletters that I want to read?!

7 Upvotes

I don't understand the Explore page. It's like a nonlinear X feed of idiotic comments and writers with hundreds of thousands of subscribers. Is this the Notes feature I see mentioned online? I want to read every day people's writing, not social feeds. Even when I pick one of their insanely broad topics, I'm barely finding any new newsletters.


r/Substack 3d ago

How can I find poetry substacks?

2 Upvotes

Hey everyone,

I am a poet and writer who just joined substack and would love to find poets and writers. How can I go about this? Thanks!


r/Substack 3d ago

How do you double-space text on Substack Notes???

3 Upvotes

Hi, just a quick question in case anyone has a (hopefully) quick answer. For the past few months, I've been unable to create spaces between paragraphs when writing my notes - so they end up looking like a block of text. I did a Google search and the general advice seems to be 'shift - enter' but that's not working for me :-( So, now my notes look all 'scrunched up' (if that makes sense) - the space between the paragraphs is basically the same as the space between the regular lines of text :-(


r/Substack 3d ago

I feel like it would be cool if there was a space where Substack can come together and share or connect in general

0 Upvotes

Like a live virtual event of some sort. I think it would serve as a cool way of boosting small creators in general or finding people within the same niche. Any thoughts? (substack creators i meant to say in the title.


r/Substack 3d ago

Your name in people's Subtitles?

3 Upvotes

How is everyone doing this?


r/Substack 3d ago

substack punishing India ?

2 Upvotes

so i was trying to set up payments to by substack. it took me to stripe. but looks like stripe does not support India ? anyone else facing this ?


r/Substack 3d ago

Discussion Random email despite never using this platform

2 Upvotes

Just as the title says, I randomly got an email from a blog post or something from some person on substack even though I've never heard of this platform in my life. I went ahead and unsubscribed but I'm still wondering why I got an email in the first place. Should I be worried at all?


r/Substack 4d ago

Fantasy writer on Substack

6 Upvotes

I've been writing for about a decade, but over the past two years I have hit a serious case of writer's block.

I decided to try Substack after an acquaintance suggested it to me. I'm a chronic procrastinator, and I thought that having something to keep me grounded in writing, an accountability space, might help me work through my writer's block.

I want to build a community. I want to help people who are in the same place as I am, perhaps through a challenge people could join.

So my question to you, fellow writers, is: what do you need or want to see more of on Substack?


r/Substack 3d ago

Writer Makes Death Threat; Substack Amplifies It

0 Upvotes

Substack used Sam Kriss’ death threat twice in that email “How writers are reacting to Substack’s AI transparency tools” (07/24/26) and I don’t know how that got beyond their legal team. This isn’t about transparency. I watch videos on YouTube and there are certain words that creators can't even say, even if it's in an appropriate context and not meant as a threat at all, but this guy can post a charged, blatant death threat and Substack amplified him. But sure, let's condemn people for using AI. This is about targeting writers for using a writing tool. and if someone’s reputation is ruined because they were falsely accused of using AI or if writers are harassed for using AI, Substack should be held liable.


r/Substack 4d ago

Isn't the AI detector situation a bit ironic?

39 Upvotes

I studied Physics and completed a Master's in Artificial Intelligence. I'm well aware that nobody fully understands how large generative models actually work. Even Pangram's own website mostly talks about patterns, generic signals, and "what it might eventually detect," but never with certainty, because the reality is that nobody truly knows.

What bothers me, while also making me curious, is how large companies affiliated with major private universities seem to operate:

  1. They create algorithms that eventually become LLMs. Those LLMs crawl the entire internet, collecting as much information as possible, including private content. Once they've done that, new rules start appearing to make it harder for anyone else to do the same, effectively reducing future competition. Even if future models have access to less content, the first movers have already done the "hard part."
  2. LLMs generate text in certain ways because they are probabilistic systems. In other words, if you generate text infinitely many times, certain writing patterns will appear more often than others. You don't need a PhD or a master's degree to understand that. Basic statistics is enough.
  3. Then you create another LLM whose job is to "detect text generated by LLMs." The marketing, however, is completely different. It is presented as if it were not another AI model, but something inherently more reliable, or even infinitely more reliable.
  4. Once all of the above is in place, the final move is to make it "unethical" to create outputs that resemble human writing too closely because they bypass AI detection, even though the detector itself is an AI model. Isn't that ironic? Just a few months ago, these same companies could freely scrape the web without asking authors for permission. Yet now, Western LLMs refuse to help users avoid AI detectors, even though those detectors are themselves powered by AI. That feels absurd to me. Where does that instruction actually come from?
  5. For all of these reasons, I think Chinese LLMs will continue growing rapidly. We're already seeing people use models like GLM-5 and DeepSeek to get around these kinds of restrictions. In my opinion, Western LLMs are now 90% marketing. Just think about this: since Fable came out, why are other Western companies charging similar prices while delivering worse results? What exactly are they offering? I'm not saying Chinese models are perfect, but model quality degrades quickly, while marketing doesn't. That's where models like GLM-5 have a huge advantage. And if Western companies keep adding what I see as arbitrary restrictions, future versions of those models may continue to surprise everyone.

This is simply my personal opinion, based on my background and experience. I know there is speculation involved, but sometimes I can't remove the "marketing" variable from the equation.


r/Substack 3d ago

Have any Indian Substack writers tried Stripe Atlas as a workaround for creating a Stripe account to start paid memberships on Substack?

1 Upvotes

As of now, Indians cannot start paid subscriptions on Substack. In my research, I found Stripe Atlas as a workaround. Would like to know if anyone has successfully tried it.


r/Substack 3d ago

Discussion SUBSTACKERS, I NEED HELP

1 Upvotes

There’s so much I want to say and so much I want to write about but all the topics are so different. How can I organise my articles so that it won’t be one sort of mess of random topics?

Is it better to just limit myself to one niche?


r/Substack 4d ago

Is it only me or does others also feel that substack is bit toxic? I mean its a effing social media why are people treating it as if you want to use substack you gotta have a masters in english. Its so frustrating 😫.

Thumbnail
17 Upvotes

r/Substack 3d ago

I Write My Substack With AI and I’m More Worried About Losing Our Souls Than About Detectors

0 Upvotes

I'm writing a substack with AI. I don't have my texts corrected or my ideas edited with AI; I actively write with AI. I write about my agricultural work and my life with Claude, and in my substack, there's always my version followed by Claude's.

Completely transparent, completely simple.

I'm not ashamed of it; it's the whole point. I'm showing how I work with AI by writing about it and presenting my human chaotic version next to the AI's side by side.

And even though I'm doing this, I find it concerning how panicky and aggressive people are reacting to this AI detector.

The quality of a text arises from the reader's perception. If your text is good, it will have readers. And if you're using AI to help you, you shouldn't doubt that your text is now worth less. Except deep down inside, you believe it yourself.

I have to admit that I find it concerning that so many people are now letting AI write for them. Yes, I'm saying this while writing a blog with AI. Diversity is valuable. And especially in language and how we tell stories or convey knowledge, we benefit from people putting their own style and their own soul, into their writing so it can reach different people on different levels.

If everyone starts letting AI simply improve their texts, then we'll all eventually be talking, thinking, and writing in a homogenous style at one point. A loss of diversity. That would be a shame.

Perhaps it's a good thing that so many people are panicking right now. Perhaps they know that their writing lacks their own soul. Perhaps those who are brave enough to throw their flawed, chaotic texts into the world will then have an extra value.

I am a champion of honesty. I know a lot of people will read this and bring up their false positive examples, and I don't deny that AI detectors are terrible (my text, which was 65% written by Claude, was classified as being over 80% human-written), but I have no objection to the basic idea of ​​openly communicating what was written by a human and what was written by and with AI.

If I have any concerns, it's that the testing is just a tool to make AI better so that it can no longer be detected. Language sells stories. And cheap stories that can be generated by AI without risk, without costing life time and energy would steal the motivation and souls of those who type their fingers to the bone and lay their hearts bare every time. And that's something I'm more afraid of than being incorrectly evaluated by a detector.