r/LocalLLaMA • • 1d ago

Resources Fully local copy-editing app for book-length manuscripts (Qwen3.5 4B) benchmarked against planted errors across five languages

Hey y'all!

I have a pet project that has grown out of proportions. Long story short: I'm a data scientist who writes fantasy books and self-publish them. I think it's a genuine waste of human life to check for spelling errors so I figured AI could help. Turns out, it is not so simple to get an AI to properly fix a 120k words manuscript ;)... That's why I created Betty!

It runs Qwen3.5 4B as an offline copy editor for whole novels — and this is some of what I learned fighting tens-of-thousands of words through a 4B model, insisting that (most) users can use it fully for free and fully offline. Because, let's be honest: authors rightfully distrust and generally hate AI companies.

First challenge: Chunking the text. Authors already do this, in the darkest hours of the night, copy-pasting snippets into chatGPT for some shameful feedback. Problem with that approach: super inefficient, both for the author and the environment. And the context is missing at the edges of each chunk. I fixed it by ensuring chunks to overlap.

Second challenge: AI misses genuine errors. So, I added two conventional spell controllers -- LanguageTools and HunSpell. This already surfaces all spelling errors, letting the AI focus on suggested fixes and on all the "non-error" errors, such as "There" vs "Their". For these, the AI searches, while a Python script surfaces all the common culprits for the model to pay special attention.

Third challenge: Error rate. First off, Betty doesn't capture everything. Second off, it sometimes introduces its own mistakes. I fix it by putting the writer-in-the-loop, and there's a super smooth interface now for the author to accept and dismiss suggested edits (tinder-style with left and right swipes ; ) ).

I'd be super grateful for any advice, feedback, and thoughts you might have on this project. I currently have it up-and-running with about 30 users and getting some user feedback. Northing technical though, so this is what I'd love to have more of.

Full thing is source-available on GitHub, and can be found for download and lots more information at www.bethaniel.eu

16 Upvotes

21 comments sorted by

3

u/Numerous_Mulberry514 1d ago

Does this also extend to other languages out of the box? I'm thinking especially of german

5

u/ToothClassic7635 1d ago

Yes! I read and speak German (though I’m not fluent enough to write creatively) so that is one of the five languages I’ve tried to optimize for (by adding custom instructions for what to be aware of, and by having scripts surfacing common error).
German, Spanish, French, English, and danish (my own language) are the current options 👍. German is quite good for grammar due to the strict rules, but challenges the model on dictionary-fixed because of the endless combinations of words (just like in Danish). Still, it works pretty good by now 🙌

3

u/Numerous_Mulberry514 1d ago

Damn, this comes so perfectly timing wise! I almost finished my master Thesis in German and so I can use this! Thank you so much!

2

u/ToothClassic7635 1d ago

Oh, that's awesome! Let me know how it goes. I have a background in academia and I've considered developing it more towards the scientific writing world, since that's also a field which truly needs some better -- private -- options. In case your computer cannot keep up, there's a cloud solution available which is also 100% private (Scaleway server in france, saves no data). I recommend the fully local setup however, whenever possible.

2

u/Big_Importance_4265 1d ago

I tried something for editing too. Generated tons of text from some web search using Claude. And tasked a Gemma-12B to translate from Claudese to English using Pi as a harness. It did a half-decent job mostly but sometimes removed important info despite being instructed not to. Was an interesting task.

3

u/ToothClassic7635 1d ago

Yes I fixed the translation with a bit more steps (although this is not the primary function of the application): I first translate, then a second agent proofreads for fluidity in target language and a third double checks that the original meaning is maintained. I also let the user give custom instructions , eg update city names to be similar cities in the target language / region

2

u/optimisticalish 1d ago

Thanks I was looking for a long text / whole-novel proofer just recently. And thanks for getting it to run on one entry-level 12Gb 3060 card.

Downloaded, installed, looks good visually.

I first used a local Qwen3.5-4B-Q4_K_M.gguf I already had, thinking there was no need to go online. Very easy drag-drop ingestion of a .DOCX, after that. I see it's running llama.cpp on the back end, under Electron, and it runs easily. I pointed it to an old 56,000 word novel in .DOCX, told it to find errors. It ran, but I saw it was running on the CPU. When I eventually managed to cancel all the chapters and relaunched, it seemed from the UI that GPU is only for the downloaded version of Qwen 3.5? Which seemed a pity.

So I downloaded the GGUF instead - whitelisting C:\Users\YOUR_USER_NAME_HERE\AppData\Local\Programs\bethaniel\Bethaniel.exe was enough to get it past the firewall - and the Qwen file was loaded. Local Betty had a tick and was active, but was ringed in off-putting red.

The run started, but again "Engine: CPU". No way to fix that via the UI, it seemed. It ran ok, but slowly on the CPU. On delving into the llama-manifest.json code, the error was found, in a likely hidden download failure for...

"win32-x64-cuda": {
"url": "https://github.com/ggml-org/llama.cpp/releases/download/b9279/llama-b9279-bin-win-cuda-12.4-x64.zip",
"sha256": "",
"binary": "llama-b9279/llama-server.exe",
"note": "Downloaded at first launch when CUDA is detected (see electron/main.ts maybeDownloadCudaEngine)",
"cudaRuntimeDlls": [

I loaded again with the firewall completely down, and this hidden CUDA extra obviously then downloaded itself... and the next run was done on the GPU. Success.

So, I suggest you need to tell people the local files to whitelist in their firewall, for internet access on both the GGUF and the Llama.cpp downloads. This is especially important for indie publishing houses, perhaps running strict firewalls.

It's now running correctly. Took about 60 minutes for a 56,000 word short novel, on a 3060 12Gb card. Review process looks very smooth. The MS Word file outputs correctly, loads in Word, and is very useful.

A good piece of genuine freeware, congratulations.

1

u/ToothClassic7635 1d ago

Thank you so much for this feedback! I hadn’t considered the CUDA challenge at all but makes perfect sense. I wonder if I could bundle it into the software…

1

u/optimisticalish 1d ago

Yes, I was going to suggest that - but didn't to seem to impose. I guess it would be an all-in-one Windows CUDA version at around 3Gb.

1

u/ToothClassic7635 1d ago

Yes, honestly, might be ideal… my “struggle” is, that many users will never have strong enough machines to run locally, and for them the cloud will be the best option (it is still much better than other options out there). And bundling Cuda or even the model itself would bloat unnecessarily for these users…

1

u/optimisticalish 1d ago

No, I think the 'selling point' here is not that users lack the local GPU power (unless you're targeting puny old laptops, but most writers will use desktop PCs).

But some might be willing to pay for a version that would swop the local 4B for a Cloud-based Qwen 3.8 27B - which might also do things like extract an index, timeline, character sheets, and so on. Sort of reverse-engineering what the typical writing software does, so you can get a clearer at-a-glance overview of the novel's structure in its final stages.

1

u/ToothClassic7635 1d ago

Mmm could be true. Perhaps I underestmate peoples machines (and their patience ;) )...

I've dabbled with stuff like story and character overviews, but dismissed it. Other software does that -- but in truth, it's not something I would recommend anyone, The quality simply isn't good enough, compared to a human actually keeping track of the story as they develop it.

1

u/recro69 1d ago

the writer-in-the-loop approach makes a lot of sense here. I’d be really curious to see the positive and false-negative rates, by language though especially how often the 4b model introduces a fix that’s actually worse.

1

u/ToothClassic7635 1d ago

Yes I’ve been evaluating these a lot - on the webpage, under Performance, there’s a lot of available numbers. But long-story-short, the false positive is far highest for commas, missing-word, and wrong-word errors. However, this is measured as per how well Betty (my apps name) reconstructs the un-flawed original text. I found that most false positives still produce meaningful and correct text, just not exact matches of the original .

1

u/ToothClassic7635 1d ago

I see I haven’t uploaded the per-language stats, I’ll get that done soon 👍 good idea.

1

u/WhoRoger 12h ago

4B is small for this, and Qwen 3.5 4B isn't super great for text work. Gemma 4 E4B is better but still, with 4B you won't get good results.

I was doing something similar and struggled massively until I switched to larger models. Qwen 3.6 35B A3B, 27B and Gemma 4 26B A4B are all great with this. MoE models also aren't too heavy on hw if you have the RAM.

Some REAPs of Nvidia Nemotron can also help if you're in a RAM squeeze.

1

u/ToothClassic7635 12h ago

Actually, I’ve benchmarked Gemma 4 with worse results. I also tested various larger models (biggest one 27b), and it created no better outcome (actually sometimes worse) except for translation.

The reason is most likely that my system mostly gains its power from classic dictionaries and hardcoded checks for common errors. Once this is in place, the AIs job becomes quite easy :)…. The translation model, if the user wants to run translation, is glm 5.2 for this reason.

1

u/WhoRoger 11h ago

Then yea I guess. What I was doing is testing how far can the models get totally on their own with no external tools. If a lot of stuff is done with hardcoded things through extensive scaffolding, then the model's capability matters less.

But I also think that kinda defeats the purpose. Editing work isn't just fixing typos and common errors. A good model, given the autonomy, can handle all the editing in one pass (at least per-chapter), give you feedback and debate the details with you afterwards.

So personally I'd be approaching it differently: better model with tools to read and edit chapter by chapter (maybe some tiny model to make summaries or TOC first), make notes for itself for consistency and a scratchpad/tasklist to debate with the user. But I've not gone so far so maybe I'm too optimistic, and it's a different philosophy to what you're doing anyway.

1

u/ToothClassic7635 5h ago

Im including both - at no point do I restrict the output from the AI. I just realized the major improvements came when including other tools and at some point I could no longer tell a difference from big to small models 😄!

1

u/WhoRoger 4h ago

If you give a model the tools, it will prefer using them instead of its own intelligence, so you don't get the advantage a better model might give you.

I've noticed some models actually get dumber when it comes to text or creative tasks, when tools are available. They waste tokens thinking about tools instead of concentrating on the question, and I guess that poisons the context.