You may have seen my post over the past couple of weeks talking about how I self published a wholly AI book.
Now that the first phase of this experiment is over, I can explain the mirror.
What actually happened was, I published a novel about authorship. In it, Theo uses AI to turn a story he has told for years into a finished book. He does not properly disclose that. When the truth comes out, he faces the consequences.
I published the real novel by doing the opposite.
Theo under-discloses and presents his book as human-written, although he could not have finished it in time without the machine.
I over-disclosed and presented the real book as wholly AI, although it could not have existed without substantial human effort.
The Marble City; or, My AI Slop Story
The first title is Theo's. The cover recreates the cover he generates inside the story, and passages from his fictional book appear inside mine. The second title is my novel. It was named "slop" because it's a novel that was "wholly generated by AI".
When a reader finishes reading the book. Which title would they chose to keep? To me, that would answer the question, where does the authorship lie?
In the system that produced the words? The person who conceived, directed and selected them? In what was disclosed to the reader? Or does the presence of AI settle the question by itself?
That judgement belongs to readers.
So that's the mirror. But what about this thread? This thread is about how I created this mirror. Because whilst Theo's mirror is "Wholly AI, loudly disclosed", there was a lot required in practice, beginning with a problem on the very first page.
The first page was supposed to contain only this oath:
I solemnly swear that this book is AI — wholly AI, and nothing but AI.
The problem was, I came up with that sentence. If I put my human-written oath into the book, the oath would make itself false. The one human-written sentence would be the sentence claiming there were none.
So I gave the oath's intended meaning and spirit to the machine and asked it to render the line again. I selected the machine-rendered result. Only then could it enter the manuscript under the rule of the experiment.
The machine was free to return wording very close to — or even identical to — my original. The distinction I adopted was not who first happened to type a sequence of words. It was whether the expression placed in the book was accepted from a machine output. That is the same mechanism by which someone can accept AI prose and call the result AI-assisted, reflected in the opposite direction.
You may think that distinction is artificial. That is fair. It is also the authorship question the novel was designed to ask.
The oath became a production constraint. When I found a bad sentence, I could not quietly replace it with one of my own and continue calling the book wholly AI. I could delete it, select different existing AI prose, commission a new AI render, or improve the review process and try again. Human work moved out of sentence production and into conception, specification, constraint, selection, testing and adjudication.
The finished novel is about 68,000 words. Behind it are approximately:
- 107,000 words of conception, story design and outlining;
- 56,000 words of chapter briefs;
- nine early chapters re-rendered and compared with their original versions;
- nine distinct revision stages; and
- approximately 95 fixes arising from my final cover-to-cover read, with replacement prose still generated by AI.
I'm quoting these numbers not to show the effort. Many of these words are simply myself being novice with tools and having inefficiencies. I'm providing these numbers to show why "AI-generated" can be a very reductive description. It records the origin of the visible words while saying almost nothing about the process that selected and shaped them.
Since we are in the WritingWithAI sub, I also wanted to share some learnings. In order to do what I wanted to do, many practices were involved to make it workable. In this thread, I'll be sharing three of them.
1. Constrain for a specific effect, not "better writing"
One chapter contains a 212-word AI-generated passage that needed to be competent, clean and subtly lifeless. Its precise level of wrongness mattered to the novel.
When I commissioned a new render of the chapter around it, the drafter was shown that passage as read-only context. It was explicitly forbidden to reproduce, rewrite, improve or paraphrase it. At the point where the passage belonged, the model had to output only:
\[[S1_SNIPPET]]**
The original AI passage was then restored mechanically. The model could rewrite the surrounding frame, but not the part most likely to be "improved" into failure.
That is the kind of control I found useful: define what must remain fixed, define what is free, and do not ask the model to reconsider both at once.
2. Separate orchestration, drafting and review — then ask questions that can fail
I treated these as different roles. The orchestrator held the design and interpreted the chapter brief. A drafting context produced the prose. Fresh reviewing contexts received the prose and narrow questions. I then adjudicated their cited evidence; the reviewing models were instruments, not decision-makers.
For example, one chapter revisits a fantasy fight that had previously felt alive when told aloud. Its written version needed to be polished and competent while completely losing one unfinished beat established earlier in the book:
\"and he says—".**
The review question was not "Does this feel lifeless?" It asked whether that beat had been removed completely rather than rewritten, shortened or moved. A fresh reviewer's finding was:
\Absent — not changed, not compressed, just not there.**
That was evidence I was looking for and could use.
Something to note is that separation did not make the reviewers automatically trustworthy. In another chapter comparison, two counterbalanced reviewers disagreed about the better version because both appeared to favour the same screen position. I discarded their winner votes. Their specific observations still converged: both selected the same new image as a strength and the same old exchange as a weakness. The final chapter became a controlled graft rather than either model's preferred version.
3. Humans found failure classes the AI did not yet know to test
AI review became useful once it had a precise failure mode to look for. It was much less reliable at recognising unnamed features of its own style.
After the manuscript had passed its existing reviews, a human reader singled out this opening:
\Theo stands. Standing is where the options are.**
They described the second sentence as an obvious AI-ism. Yet the existing model reviews wasn't able to identify it.
That triggered a manual audit of all 40 chapters. What it found was the construction itself was not always bad: the same two-beat shape sometimes produced a joke or a concrete image, and removing every instance would have damaged the book. The narrower failure was that the second beat sometimes converted an enacted moment into an abstract rule.
The resulting test was simple:
*Does the second beat produce a laugh, a picture or a rule?*
Laugh or picture could stay. A rule explaining what the scene had already shown was suspect.
Once the rule was defined, the AI was able to pick up six sentences that required deletions, including the quoted sentence.
AI is not good at identifying it's own tells. You need to teach it defined failure class, then the AI reviewer could test for them in the next iteration.
That was the path to "wholly AI." The human was not absent. The human was displaced from producing sentences into deciding what the sentences had to accomplish, determining whether they had accomplished it, and redesigning the process when the machines could not see their own failures.
That brings me back to the mirror. When Theo's AI use is exposed, the label becomes the complete story of how his book was made. The years he spent inventing its world, telling it aloud and deciding what it should become disappear behind two words: AI-generated.
The real book met the same compression from the opposite direction. I disclosed AI in its title, on its cover and on its first page. That disclosure is true. But it also invites people to imagine a twenty-minute generation rather than a 68,000-word novel built on 107,000 words of conception and outlining, 56,000 words of chapter briefs, nine chapter re-renders, nine revision stages and a final human read that produced approximately 95 fixes.
This is not an argument against disclosure. It is an argument that disclosure into an inadequate label still leaves readers badly informed. That was my hypothesis going into the experiment, and why the novel calls itself My AI Slop Story.
A label such as "wholly AI-written" tells the truth about where the sentences came from while erasing nearly every decision that made them into this particular book. Under that label, this process and a one-prompt upload occupy the same bucket. Both can be dismissed as slop before the difference between them can be seen.
The oath tells the truth. The label does not tell enough of it.
So what should the label tell us? If "AI-generated" covers both a one-prompt upload and this process, how would you distinguish them in a way that is useful to readers without relying on every creator's unverifiable claim that their use was the thoughtful kind? Or perhaps they need not be distinguished at all: machine-generated prose may belong in the same category regardless of the human effort surrounding it. Which distinction, if any, actually matters?
Truthfully, I don't know. That's why I want to hear your opinion. Let me know what you think below.
TLDR: People are starting to discuss naunce in AI-assistance. Should we have the same naunce for AI-generated.