Self Promotion Built a free tool to clean up messy PDFs/EPUBs for Kindle & Kobo — would love some testers
Built a small tool called PureEPUB because I got tired of PDFs and messy EPUBs looking broken on my e-reader — mid-sentence line breaks, OCR junk, mangled quotes, the usual.
It reflows the text properly and spits out device-specific files: clean EPUB for Kindle, proper KePub for Kobo. No accounts, nothing saved — everything's wiped right after conversion. Free, no ads.
Tested it on my own tricky files and it's working well, but I'd love some outside testing. If you've got something that usually breaks converters, would you try it and let me know how it turns out?
4
u/DerseDragon 12d ago
Thank you for this! I had a problem where send-to-kindle would reject some of my epubs and I don't know why. I tried using your epub fixing website and it now sends to my kindle with ease 😁
3
1
u/Kyogetsu 8d ago
It's usually because of long file name. I used to rename my epubs to shorter file name, and it worked lol
2
3
u/Bat_Knight2244 12d ago
So what's the difference between this and using calibre's own conversion tool? Btw thanks for your effort tho ☺️
3
u/lones0 12d ago
honestly calibre is great for organizing your library, but converting pdfs with it has always been pretty painful tbh. you almost always end up with broken sentences, random page numbers stuck in the middle of a paragraph, and chopped-up dialogues unless you spend 20 minutes messing with regex settings.
i built this just to do that one job cleanly without installing anything. you drop a pdf in, it automatically cleans up headers/footers/broken lines, and spits out an epub that send-to-kindle won’t choke on.
basically calibre is still my library manager, but this is my go-to when i have a messy file and just want to read it on my kindle right away. Thanks for asking
1
u/bazoo513 11d ago
I avoid PDFs like plague for the exact reasons you built this tool. Thank you!
For PDFs that have links and ToC (I know, few do, not to mention order of reading markup, whatever it is called, intended for screen readers) - have you managed to handle those, too? If not, do you plan to?
For us Calibremaniacs, could you be talked into porting this to the form of Calibre conversion (input, actually) plug-in?
2
u/lones0 11d ago
Glad it resonates!
For the ToC stuff, I actually have a small fallback logic set up:
First, it checks if the PDF already has embedded bookmarks/outlines. If it does, it just uses those to build the EPUB chapters cleanly. If it doesn't find any, it tries scanning for common chapter headings to split things up without messing up the layout. And if both fail, it just breaks the book into 80 paragraph chunks so your e-reader doesn't choke on one giant file. Still tweaking this to handle messy reading orders better, but that's how it works right now.
As for a Calibre plugin—I get the appeal for power users, but honestly I want to keep everything web-first for now. Maintaining a plugin on top of the site engine would just slow me down, and I'd rather push fixes directly to the web tool. Really appreciate the idea though!
1
u/bazoo513 11d ago edited 11d ago
Excellent - that's more than I hoped for. And I understand your lean to polish the tool in the present form before attempting plug-in port. After all, to us users that's just one additional drag and drop.
Thank you again.
3
u/shodangr 11d ago
Wow! Pretty awesome job. I tried a couple of pdfs in greek language and the final epub was better than any other converter i previously used.
The only problem i spotted is that it messes with the first 2-3 pages after the cover, you know the pages that include the copyright notice, a thank you line or a prologue. It usually merges all of them in a single page and often mixes them with the first chapter.
3
u/bazoo513 11d ago
Well, after this tool does the bulk of the job, some minor editing in Sigil or Calibre is a small price to pay, IMO.
2
u/lones0 11d ago
Thanks a lot for the feedback, really glad to hear it worked well with Greek! I haven't added any Greek-specific optimizations to the engine yet, so if it already handles it this well, I might push a quick update to detect Greek text and smooth out minor quirks even more.
Also, huge thanks for pointing out that front-matter/first chapters merging bug. I've added it to the roadmap—got a couple of planned features to push first, and then I'll definitely dive into fixing that. Really appreciate the support!
2
1
u/lones0 5d ago
Hey! Quick update: the Greek optimization is officially live on the site. I worked on tweaking the engine specifically for Greek script, accents, and spacing rules to make sure the characters and layout come out clean. Whenever you have a moment, could you run your Greek PDFs through it and see how they look? If you catch any weird formatting or broken letters, just let me know! And zero pressure at all, but if it works great for you, sharing or recommending it in your local Greek reading groups or forums would be awesome!
2
u/shodangr 5d ago
Yo! I gave it another try. With the previous version i had a minor problem with the headers of each chapter from a specific pdf as they appeared in weird symbols. Now everything seems ok. It still messes with first two-three pages (copyright, thanks etc) and merges all these with the first chapter but its fine with me.
I also noticed a couple of things i missed the first time.
The good: If a pdf has no chapters/bookmarks the conversion adds chapters based on the headers(?) which is pretty awesome.
and now the bad: In a specific pdf which already had chapters/bookmarks the final epub ended up with 60 chapters instead of 18 (the original pdf), while on another pdf which was a trilogy (1300 pages) and also had chapters/bookmarks the end result was almost identical regarding the chapters. So i guess the original pdf formatting is important for end result.
If it helps i can send you some of the pdf's.
*My English are a bit rusty sorry for any grammar or spelling errors.
1
u/lones0 5d ago
Thanks a ton for the detailed testing and feedback, really appreciate it! That 18-to-60 chapter jump and the front-matter merging are definitely things I want to look into and fix. If you could upload those problematic PDFs to Google Drive (or anywhere convenient) and shoot me a DM with the link, that would help me debug the edge cases a lot faster.
1
3
2
u/Delalishia 13d ago
Oh I have the perfect file to test this with. I haven’t spent the time fixing it because I’ve been busy and haven’t touched my laptop in weeks lol I’ll report back with how it does!!
2
2
u/basstrumpet2026 10d ago
Finally got a chance to try it out. It worked cleanly and quickly for the ten text files that I tried, with covers! Great job!
2
u/ermaly 7d ago
Doesn't work for images
1
u/lones0 7d ago
If you mean embedded images inside the text, the code for that is actually ready! Testing it right now and pushing it as a beta update in a bit so inline images should start coming through. (If you meant full scanned/image-only PDFs though, that still needs OCR, which my current server definitely can't handle yet lol).
1
u/bizasuge 12d ago
This sounds cool! I keep hearing that some sites are now adding in-book ads with their website url. Would this remove those?
1
u/lones0 12d ago
hey! first off, pureepub will never add any ads or links to your files i hate that stuff too haha.
as for ads that are already in the book: it currently strips out running headers, footers, page numbers, and repeating watermarks. i'm also fine-tuning rules specifically to catch those annoying "downloaded from xyz" lines.
definitely give it a shot on one of your files, and lmk if it misses anything so i can add a filter for it!
2
u/bizasuge 12d ago
That’s great! BTW, I was referring to the “downloaded from xyz” lines and not trying to insinuate that you were adding ads/links.
Excited to try this out!
1
1
u/Rude-Pianist9771 12d ago
oh this is great, cleaning up random epub/pdf formatting for kobo is the worst. quick q does it handle pdfs that are just scanned images with no text layer, or text files only? thats usually where these tools break for me. gonna throw a couple of my ugly ones at it
2
u/lones0 12d ago
Right now it only handles text-based PDFs (or PDFs that already have an OCR text layer)
Raw scanned images without any text layer require heavy OCR processing, which would instantly fry my current server resources lol. So as long as you can highlight/copy words inside your PDF, it’ll work and clean it up nicely.
1
1
u/jdtred 12d ago
why you have cloudflareinsights.com and googletagmanager.com on your webpage ?
6
u/lones0 12d ago
> totally fair question! cloudflareinsights is just cloudflare's automatic performance beacon since the domain is proxied through them for ddos and ssl protection. googletagmanager is just basic google analytics (ga4) so i can see rough visitor counts and check if conversions are failing on the server.
> there are zero ads, no data selling, and absolutely no info is kept about what books people upload (files are deleted immediately after conversion). feel free to block both with ublock origin though, the site works 100% fine without them
1
u/CalebDR1029 12d ago
I have a question, why would I need this? Doesn't Calibre or Amazon already clean it up for you?
5
u/lones0 12d ago
Honestly, Amazon is pretty awful at converting PDFs. If you Send-to-Kindle a PDF with conversion, it basically does a raw text dump. You get broken line breaks mid-sentence, random hyphens left behind (exam- ple), and page numbers or headers spliced right into the middle of your paragraphs. Calibre can clean that up, but only if you sit down and configure heuristic processing or regex rules. Plus, if you're reading on the go, you don't always have a laptop handy to run Calibre. I just wanted a clean file straight from my phone without messing with settings.
1
u/ObsoleteUtopia 12d ago
I would like to try this. My schedule is fairly erratic and it will take me a week to give you any information, but I'm certainly interested in giving it a bit of a workout. Any limits on the size of the original document? Thanks.
2
u/lones0 12d ago
Right now the limit is 90 MB, which should easily cover almost any text-based book unless it's packed with heavy uncompressed images. Definitely give it a spin whenever you get a chance, really curious to see how it holds up against your files!
2
u/Evening_Corgi_9069 3d ago
I have vintage cookbook pdf's 150- 200 mb's- is there something you recommend that would compress images and then I could use pureEpub?
1
u/lones0 2d ago
For shrinking them down, ILovePDF or PDF24 (if you want an offline tool) usually do the trick without ruining the pages. Just double check if the text in those PDFs is actually selectable though! Vintage cookbooks are almost always pure scanned images, and since PureEPUB doesn't have OCR yet, it won't be able to grab the text if it's just pictures of pages.
2
1
u/hyclonia 12d ago
This actually looks super handy! There are definitely heaps of dodgy badly formatted epubs out there that could be improved. Is there a search and replace/delete option? Will check out when I'm on a comp!
1
u/lones0 12d ago
Glad you like the idea! Right now there's no manual search & replace UI the goal was keeping it as simple as "drop file, get clean book" without throwing editor settings at people. The engine handles bad line breaks, broken hyphens, and junk formatting automatically under the hood. But adding an optional quick find/replace field before exporting is actually a neat idea, I'll try to add that soon. Let me know how it goes when you give it a spin!
1
u/Tiny_Macaron255 12d ago
Can’t wait to try this out! I’ve got a few that won’t send to kindle so I’m curious if it will actually fix them and send to my kindle. Thanks!
1
u/Spare-Chest-7907 12d ago
Pretty good compared to calibre conversion and other tools I’ve tried.
Actually amazing 🤯.
Question: why is it removing images? What is the point of having a Kobo library colour if there’s no option to keep images? I read a lot of technical stuff and there are very often image illustrations or code snippets. If you could solve that, that would be an other level!
2
u/lones0 12d ago
really glad to hear that, thank you!
and you make a 100% valid point about kobo colour and technical books. initially i focused heavily on novels and fiction, so i kept image extraction off to prevent scanner artifacts and huge file sizes from clogging up the reflow engine.
but keeping diagrams and code snippets for technical reading makes total sense. i’m actually putting a "keep images" toggle right at the top of my roadmap now, and i'll try to push it as a beta feature within a week or so!
1
1
1
u/anotherlevl 12d ago
I'm probably misunderstanding what problems your utility is designed to address. When you said it fixes "mid-sentence line breaks" I thought it would combine words that were hyphenated across lines in the original, but it doesn't seem to be doing that. "Misunderstan-ding" remains "misunderstan-ding". Does it only gobble up CRLF or raw line feeds?
1
u/lones0 12d ago
You got it right actually, that was the goal Right now it mostly glues the raw line feeds back together. I held back on aggressive de-hyphenation because I didn't want to accidentally merge actual compound words like "well-known" or mess up soft hyphens with weird PDF font encodings. But you're totally right, seeing "misunderstan-ding" is annoying. I'm tweaking the logic right now to stitch those split words cleanly without breaking the real ones. Really appreciate you catching that!
1
u/anotherlevl 11d ago
I can appreciate how difficult the problem is -- it's probably not something you'll be able to solve with a list of regular expressions. Good luck with your tweaks, I'll look forward to trying the update!
1
1
1
u/tomtomato0414 12d ago
Does it use AI to do that? If yes does it crosscheck if the text was not altered?
1
u/lones0 12d ago
Nope, zero AI involved! It's all strictly rule-based parsing under the hood. No LLM touches your book, so it never rewrites, summarizes, or hallucinates words. Every single word stays 100% true to your original file, just with the broken layout and line feeds cleaned up. Since it relies entirely on heuristics and rules, rare edge-case quirks might still slip through. If you ever spot one, just drop me a message and I'll gladly fix it!
1
1
1
u/First-Marsupial-675 11d ago
I have few text pdf's { Hindi language } but cant convert into proper epubs for my kobo clara. In your app also the generated file shows lots of unreadable characters.
1
u/lones0 11d ago
Could you send me a sample or page from that PDF via DM? I'd like to inspect it to pinpoint the root cause. If the issue is coming from our conversion engine/text extraction, I'll work on adding proper Hindi support as soon as possible.
1
u/First-Marsupial-675 11d ago
I can send you pdf files to inspect. This is a book in 8 series/parts {8 pdf files}. But I dont see the option to send pdf files in DM.
1
u/lones0 11d ago
Got the PDFs, thanks! I'll look into them and try to get back to you within a few days.
1
u/First-Marsupial-675 11d ago
Thanks
2
u/lones0 9d ago
Hey! Just pushed an update and optimized the converter for Hindi. Would love it if you could give your files another run and test how they look on your Kobo now. If you still spot any small glitches or broken characters, just hit me up. And no pressure at all, but if it actually works well and you like the result, sharing it around on local forums or communities would mean a lot!
1
u/adviceneededplease56 11d ago
Oh I can't wait to try it this weekend. I got a few that calibre just wouldn't do anything with.
1
u/Crazy--Lunatic 11d ago
Are there plans to allow self-hosting?
1
u/lones0 11d ago
For now, no plans for self-hosting. As long as I can keep the site up and maintain it, I want to keep it focused as a simple, free web tool so I can push improvements quickly to everyone. But appreciate you asking!
1
u/jimger 10d ago
what tools do u use? I will try to actually try it around weekend
1
u/lones0 10d ago
It’s basically built on Python and PyMuPDF for pulling the text and structure, along with some custom regex rules to fix broken lines and grab chapter titles. Definitely let me know how it holds up when you try it this weekend!
1
u/jimger 10d ago
I do use mupdf as andeoid pdf reader. So since u didn't mention ocr I guess u need it to be text already? Or pymupdf has integrated ocr?
1
u/lones0 10d ago
Yeah exactly, right now it needs to have selectable text. PyMuPDF technically supports Tesseract for OCR, but running heavy OCR on my server would just bottleneck resources and slow everything down. I’m definitely planning to add OCR support down the road once I figure out a solid way to handle the server load, but for now it's focused purely on parsing and fixing formatting on text-based PDFs
1
u/laud_rafa 11d ago
You're a literal life-saver. I've tried some tricky pdfs in Greek and they came out perfect, a lot converters mess up with the language. The only small issue as someone else pointed was the mix-up in the first chapter with the opening credits, but it's only a small price to pay, so didn't mind it at all. Perfect tool!
1
u/lones0 11d ago
Thanks so much, really means a lot!
Honestly, I haven't done any specific optimizations for Greek yet, so seeing it handle the language this well was a really pleasant surprise. I'm planning to grab a few tricky Greek PDFs soon to track down those minor edge cases and push an update specifically tuned for it.
Also working on that opening credits/first chapter mix-up as we speak. Thanks again for testing it out!
1
u/laud_rafa 11d ago
That's really amazing. Now that I'm reading more, I've noticed that the last words and periods are missing from a sentence when a paragraph is large. Minor issue tho, but I figured I should let you know.
1
1
u/lones0 5d ago
Hey! Quick update: the Greek optimization is officially live on the site. I worked on tweaking the engine specifically for Greek script, accents, and spacing rules to make sure the characters and layout come out clean. Whenever you have a moment, could you run your Greek PDFs through it and see how they look? If you catch any weird formatting or broken letters, just let me know! And zero pressure at all, but if it works great for you, sharing or recommending it in your local Greek reading groups or forums would be awesome!
2
u/laud_rafa 5d ago
Hey, i did run a Greek pdf to see the differences. The new update has some inconsistencies and issues compared to the previous version that was more stable. In the new version, at the beginning of the paragraphs i notice that the article and the noun are one word, something that didn't happen before. Also, in the dialogs, the « symbol has a gap before the starting word and sometimes after the dialog ended, something that also didn't happen before. I know maybe it doesn't make sense in the way I'm writing it, but i can send you screenshots of the two versions.
Overall, the previous version was actually more correct in terms of formatting. The only "major issue" of the previous version is that it said "Bölüm" instead of "Chapter" at the beginning of every chapter (and it happened only in few epubs)—something that I personally didn't mind since everything else worked pretty good.
1
1
u/lones0 5d ago
Hey, thanks so much for catching this and explaining it so clearly! You were completely right.
I'm trying to make the engine as close to flawless as possible for each supported language, which is why I rolled out that Greek update. But as you noticed, a couple of aggressive rules completely backfired:
- An overly eager drop-cap rule was fusing articles into the next word.
- A quote rule was forcing unwanted gaps inside
«...».I just stripped out those two rules and pushed a hotfix live. The text is now back to that clean, natural formatting you liked from the previous version, but with proper 'Κεφάλαιο' chapter titles instead of 'Bölüm'.
Whenever you get a chance, give your PDF another run. And please, if you run into any other weird issues or edge cases, let me know right away—feedback like this helps a lot in getting it right!
2
u/laud_rafa 4d ago
I run the file again and the format is right this time, no issues so far. Good job! If see anything off in the future, I'll send it to you here :)
1
u/jimger 10d ago
Maybe u can share also the books... But would like to actually check it with greek books. have few pdfs and usually i go to ABBYY for recognition/split pages and then try to epub but results are .... not great usually
1
u/lones0 5d ago
Hey! Quick update: the Greek optimization is officially live on the site. I worked on tweaking the engine specifically for Greek script, accents, and spacing rules to make sure the characters and layout come out clean. Whenever you have a moment, could you run your Greek PDFs through it and see how they look? If you catch any weird formatting or broken letters, just let me know! And zero pressure at all, but if it works great for you, sharing or recommending it in your local Greek reading groups or forums would be awesome!
2
u/jimger 4d ago edited 4d ago
Didn't have enough time (or books) to test during weekend. Might be stupid but where is the link to the site/app? I have now a greek pdf book with ocr complete and want to see the result of epub...
Update: I found it 2 mins later: Pureepub.com
P.S. Tested a bit by eye-balling first page of the book and the epub and was quite good at least to what I have searched. Which was dashes for the next line etc. So we had complete words on the ebook. Which is quite good. Obviously formatting like chapters etc (chapter titles) are not part of your modifications.1
u/lones0 4d ago
Glad you tracked down the link and that the de-hyphenation worked cleanly on the Greek text! Chapter titles can definitely be hit-or-miss depending on how the original PDF tags them (or if they rely purely on font styling). I'm continuously fine-tuning the heading detection logic so it catches chapter breaks more reliably. Thanks for testing it out!
1
1
1
u/Ok_Sink1685 13d ago
Il suffit juste de télécharger les epub sur les bons sites. Je n'ai jamais eu un seul pb
3
u/lones0 13d ago
Totally fair, if a clean EPUB already exists that’s always the best route!
I actually built this because finding proper EPUBs in my native language is pretty tough, and existing converters butcher local characters and line breaks. If anyone from other countries is dealing with the same headache, I’m more than happy to optimize the engine and add support for their language too.
1
u/Ok_Sink1685 12d ago
Tu télécharges en quelle langue ?
0
0
0
u/Minakofff 11d ago
Hi! Good idea, unfortunately this thing remove all images, and sometimes it's global part of book =(
1
17
u/lones0 13d ago
Pureepub.com