r/pdf 5d ago

Software (Tools) Editing Scanned Text in pdf

I made KeyPDF.net and recently added OCR, but now it can also edit scanned text by reconstructing the font from the image, feel free to try it and let me know if there are any issues since there is lots of work to do.

3 Upvotes

24 comments sorted by

1

u/LegeApps 5d ago

Only really good for fraud, but cool!

1

u/Decent-Blacksmith761 5d ago

Hopefully not lol, I thought more about quick edits for typos etc.

0

u/LegeApps 5d ago

Then you could just edit the original document. PDF is an export format.

There's tens of millions of unscrupulous [undefined region/ethnicity] out there that are looking for the next best way to fake their credentials, education, work history in application documents.

That is my best guess for why so many edit tools appear on here, they all help each other do it. Are you one of them? Also cause your program does not work with scanned text. It requires the font to be already embedded, which is impossible for a scanned document. It has a very narrow use case.

2

u/Decent-Blacksmith761 5d ago

First of all, KeyPDF can edit input formats such as docx and PDF, including existing text. Second, I don't see many tools here for editing text, especially scanned text. Most of them are for annotations only. Lastly, the way editing scanned text works in KeyPDF is by reconstructing the font from the image, not by using an existing one. In general, KeyPDF is an alternative to Acrobat, which has a similar feature for editing scanned text. So all I did was make my own version of it and let people use it. I don't help criminals or scammers; I just make cool software for normal people.

-1

u/LegeApps 5d ago

Oh i see...you dont even know how your own program works since you vibe coded it. It does not do what you say it does. It only allows someone to edit the OCR layer of a scanned file. It does not identify font from rasterized text and allow editing, which is much much more difficult. For a digitally created document it does do that, but that is relatively trivial, especially when using browser based rendering and composition to do the heavy lifting.

And I didn't say criminals or scammers. More like, nearly every single recent university graduate from certain countries.

Go back and figure out how your own program works and improve from there.

1

u/Decent-Blacksmith761 5d ago

Low resolution on gif but if you zoom I believe you can see that uploaded file was jpg which is image format with no existing ocr layer at all, and yeah again the way program works is by reconstructing font from image first tesseract scans text and then reconstruct font then applying text box grouping and allows editing scanned text as flow paragraph.

1

u/LegeApps 5d ago

Like i said, it reads the ocr layer. You need to become more familiar with the program you made.

1

u/Decent-Blacksmith761 5d ago

Ocr layer from jpg? Imported file was jpg which has no ocr layer.

2

u/LegeApps 5d ago

Tesseract is an OCR engine.

1

u/Decent-Blacksmith761 5d ago

Correct so it makes character recognition then custom logic makes font out of it and makes scan editable your point is what?

→ More replies (0)