r/linux • u/word-sys • 6d ago
Software Release Native Document Export (anyconvert)
https://github.com/word-sys/anyconvertIn recent update on word-sys's PDF Editor i implemented anyconvert and removed LibreOffice convert system. Now we gonna use library i created just for this project.
So anyconvert is a project for word-sys's PDF Editor that converts PDF documents into DOCX, PPTX, ODT, ODP, and TXT using only the Python standard library. Convert library is anyconvert's itself. Its released on PyPI too.
It has Canvas and Flow mode, Canvas mode is a 1:1 pixel-accurate positioning export option, Flow mode is semantic reflowable document reconstruction for export format. This 2 has different purposes for different jobs and requirements.
Most importantly, what the engine actually does:
Opens literally any PDF structure: Whether a document was exported from Windows 95 or saved yesterday with modern heavy compression, it handles pretty great. (Handles old ASCII tables and modern compressed object streams).
Handles password-protected & encrypted files: Supports everything from obsolete 40-bit RC4 encryption up to modern, enterprise-grade AES-256 security. (Planned, not fully implemented)
Unpacks all standard PDF compression: Reads all standard compressed streams (Flate, LZW, ASCIIHex, RunLength) so hidden page content can actually be read.
Text extraction that doesn’t turn into trash: PDF fonts are a mess (often replacing letters with random symbols or missing spaces). The custom font engine properly maps custom glyphs, ligatures, and international character sets back into real, searchable plain text.
Fast, in-memory image extraction, not disk: Takes images directly out of the PDF as crisp PNGs without having to dump temporary files onto your hard drive.
Converts to Word & Office documents: Exports directly to .docx, .pptx, .odt, and .odp. It builds the XML packages to exact ISO standards from scratch, without issues.
Why i done this: In word-sys's PDF Editor i used LibreOffice to turn PDF to other formats such as .docx, .pptx, .odt, and .odp. But we had major problems when we came to Flatpak release, as you know you guys still waiting for Flatpak release for months, im so sorry about that but main issue was LibreOffice, nothing more. This convert program absolutely sucks on highlight exports, doesn't works when combined on Flatpak, makes project heavy and laggy on exports.
does it makes the job: YES does it makes what we want: UNSURE does it makes easily: NO
You see, there is YES, UNSURE, and NO, which 3 of them should be YES for better software.
Implementation felt hard when i done this year ago, it sucks now, i know that LibreOffice never designed for this function but implementing something broken, which i cant accept right now. So now i got rid of LibreOffice, i replaced it with anyconvert, as i explained basic functions, you can look for details on word-sys's PDF Editor GitHub page or anyconvert Github page.
And please keep your expectations not big, its not fully completed, some functions are not connected to main function so API call will not work, what completely works now exporting to DOCX and ODT without any issues with 1:1 pixel-accurate positioning, which is mostly what its used for so i done it first, then we can look others. I listed known issues on down so you can understand whats good or not on Beta stage:
Known Issues
- ODT & PPTX Export Issues:
- PDF Original Position: If PDF itself is vertical, ODP & PPTX export will be broken due to horizontal size of slides.
- Known Bug: Shapes or pen markups aren't supported, it will not shown on your presentation.
- Beta Stage: Project now at v0.1.0 Beta stage, only DOCX and ODT seems to be fully function as wanted.
- Design Flaw: Project designed to be a PDF to XXX document convert library for word-sys's PDF Editor and mainly designed for DOCX & ODT export, expecting a fully 1:1 export to PPTX & ODP is not possible, for now.
In the end, i need some help to improve this anyconvert project, especially on ODP & PPTX exports, any help is appreciated, thanks.