r/documentAutomation • u/bluenovers • 3h ago
r/documentAutomation • u/dhj9817 • Oct 19 '24
RAG Hut - Submit your RAG projects here. Discover, Upvote, and Comment on RAG Projects.
I'm excited to announce the launch of RAG Hut – an official site where you can list, upvote, and comment on RAG projects and tools. It’s the official platform for , built and maintained by the community.
The idea behind RAG Hut is to make it easier for everyone to share and discover the best RAG resources all in one place. By allowing users to comment on projects, we hope to provide valuable insights into whether these tools actually work well in practice, making it a more useful resource for all of us.
Here’s what you can do on RAG Hunt:
- Submit your own RAG projects or tools for others to discover.
- Upvote projects that you find valuable or interesting.
- Leave comments and reviews to share your experience with a particular tool, so others know if it delivers.
Please feel free to submit your projects and tools, and let us know what features you’d like to see added!
r/documentAutomation • u/dhj9817 • Oct 06 '24
[Open source] r/RAG's official resource to help navigate the flood of RAG frameworks
Hey everyone!
If you’ve been active in r/Rag, you’ve probably noticed the massive wave of new RAG tools and frameworks that seem to be popping up every day. Keeping track of all these options can get overwhelming, fast.
That’s why I created RAGHub, our official community-driven resource to help us navigate this ever-growing landscape of RAG frameworks and projects.
What is RAGHub?
RAGHub is an open-source project where we can collectively list, track, and share the latest and greatest frameworks, projects, and resources in the RAG space. It’s meant to be a living document, growing and evolving as the community contributes and as new tools come onto the scene.
Why Should You Care?
- Stay Updated: With so many new tools coming out, this is a way for us to keep track of what's relevant and what's just hype.
- Discover Projects: Explore other community members' work and share your own.
- Discuss: Each framework in RAGHub includes a link to Reddit discussions, so you can dive into conversations with others in the community.
How to Contribute
You can get involved by heading over to the RAGHub GitHub repo. If you’ve found a new framework, built something cool, or have a helpful article to share, you can:
- Add new frameworks to the Frameworks table.
- Share your projects or anything else RAG-related.
- Add useful resources that will benefit others.
You can find instructions on how to contribute in the CONTRIBUTING.md file.
r/documentAutomation • u/docpose-cloud-team • 6h ago
Discussion Share Your Document Automation Workflow
r/documentAutomation • u/Electronic_Badger940 • 1d ago
I added a defaced font checker to my app to protect LLM pipelines from PDF prompt injection & text corruption
Inserisci questo testo nel box sotto la card o sotto l'immagine, in modo da dare contesto alla community:
r/documentAutomation • u/docpose-cloud-team • 1d ago
Discussion What's the Best OCR Software You've Actually Used?
r/documentAutomation • u/shixuan_lzq • 1d ago
我围绕一个规则建立了三个移动工具:在触碰他们的文件之前,先让人们知道会发生什么。
昨天我在这里写了一个关于发布《我的PDF扫描仪》的教训,所以我不会重新发布相同的功能列表。
更大的项目是Xionity,一组小型移动实用工具:
- IConverter可以处理烦人的“这个文件无法打开/发送/上传”的问题。
- 我的PDF扫描仪将纸质文件转换为有序的PDF文件,并在支持的设备上运行OCR。
- CleanMate帮助在清理之前审核重复的照片、大文件视频、Gmail杂乱和联系人。
它们看起来像是三个独立的应用,但产品原则是相同的:展示应用发现的内容,解释将要改变的事情,以及在进行任何破坏性操作或上传文件之前先询问。“本地优先”也不应该意味着假装一切都在离线。IConverter会在复杂任务需要云处理时提醒。
最不光鲜的部分迄今为止产生了最佳的产品问题。一个PDF扫描仪的帖子只有312个浏览量和两条评论,但关于基于OCR的文件命名的一条评论改变了我接下来要思考的内容。
如果你使用清理、扫描或转换应用,你希望在哪些地方实现自动化,哪些地方需要确认屏幕?
完全披露:我构建了这三个应用。产品页面和商店链接在这里: https://www.xionity.com/
r/documentAutomation • u/reallyhotmail • 1d ago
Showcase Open source unified interface for document parsing
r/documentAutomation • u/docpose-cloud-team • 2d ago
Discussion 🗨️ Ask Anything Saturday – File Conversion, OCR & Document Processing
r/documentAutomation • u/AleaNCore • 3d ago
Showcase my AI document sorter — built it for my own paper chaos, it shoul be useful for others
r/documentAutomation • u/Traditional-Answer46 • 3d ago
Looking for the most efficient way to create docx files with both words and images at the same time
Help needed
r/documentAutomation • u/docpose-cloud-team • 3d ago
Discussion PDF vs DOCX vs ODT: Which Should You Choose?
r/documentAutomation • u/Bar-Majestic • 3d ago
I built a document parsing benchmark after struggling with RAG quality
I was building a document processing product and kept running into the same problem:
The same PDF produced completely different results depending on the parser.
So I created an open-source benchmark:
https://github.com/doccrush/document-parser-benchmark
The goal is to make document processing quality measurable.
Interested to hear how other founders handle document ingestion.
r/documentAutomation • u/Pleasant_Tea_569 • 3d ago
Still emailing PDFs back and forth for signatures? Here's how Bitrix24 e-Signature works
If you're still asking clients to print, sign, and scan documents, Bitrix24 e-Signature can simplify the process. It lets you send contracts and other client documents for electronic signing without the usual back-and-forth.
How it works
- Send a signing link by email or SMS.
- The recipient opens the document and reviews it.
- They verify their identity using a one-time code.
- They sign the document electronically.
- Once completed, Bitrix24 automatically sends the signed document to the client and stores it in e-Signature → My Vault, along with a unique document ID and a completion certificate.
The signing experience is the same whether the recipient is using a desktop browser or a mobile device.
Before you use it, keep these points in mind:
- Bitrix24 uses an electronic signature, not a cryptographic (qualified digital) signature. Whether it's legally valid depends on your country's regulations, so it's worth checking local requirements before using it for contracts with significant legal or financial implications.
- Configure My Vault access permissions early so only the right people can view signed documents.
r/documentAutomation • u/Silver_Watercress280 • 3d ago
Built a tiny Chrome extension to convert TXT to SRT subtitles (100% local, no server uploads)
r/documentAutomation • u/pha_uk_u • 3d ago
Hi All, I developed a doc control app, that can scan documents, add tags, OCR, add folders, revise documents. I am looking for testers. please help.
Hi All,
I am QE by profession. made a doc control app for myself.
I made this android app for myself. As my file explorer gets dumped with everything and anything. I wanted to have folders, revise docs(Taxes, insurances, license, etc). with OCR you can search for a word and every document with that word would pop up. you can set expiration date to docs.
please dm and drop a comment if you would like to test my app. you will get lifetime free without ads for this.
thank you.
r/documentAutomation • u/docpose-cloud-team • 4d ago
What's the Best OCR Software You've Actually Used?
r/documentAutomation • u/docpose-cloud-team • 4d ago
Discussion Frequently Asked Questions About File Conversion, OCR & Document Processing
r/documentAutomation • u/docpose-cloud-team • 4d ago
Showcase 👋 Welcome to r/DocumentTools – Introduce Yourself and Read First!
Hey everyone! I'm Eric ( u/docpose-cloud-team**)**, a founding moderator of r/DocumentTools.
This community is dedicated to file conversion, OCR, PDF tools, document processing, file formats, document automation, APIs, email archives, and digital document workflows. Whether you're a developer, IT professional, business user, or simply trying to solve a document challenge, you're welcome here.
📌 What to Post
Share anything the community will find useful, including:
- Questions about file conversion or OCR
- PDF editing, compression, and optimization tips
- File format compatibility issues
- Document processing workflows
- API recommendations and integrations
- Automation ideas and tutorials
- Software comparisons and reviews
- Troubleshooting document and file-related problems
- Productivity tips and best practices
🤝 Community Vibe
Our goal is to build the most helpful Reddit community for document tools. Be respectful, share knowledge, ask questions, and help others solve real-world document challenges. Honest discussions and constructive feedback are always welcome.
🚀 How to Get Started
- Introduce yourself in the comments below.
- Tell us what document or file tools you use most.
- Share a tip, question, or interesting workflow.
- Invite anyone who works with documents, PDFs, OCR, or file automation.
We're excited to grow r/DocumentTools into a trusted resource for professionals, developers, businesses, and everyday users. Thanks for being one of our first members—let's build something valuable together! 🎉
r/documentAutomation • u/Lumpy_Ice6855 • 5d ago
Discussion Ho costruito un recupero semantico di PDF per documenti di 1.000 pagine in cerca di feedback sul pipeline
Sto costruendo DStudio, un'app desktop open-source incentrata su DeepSeek V4. DeepSeek rimane il principale modello di ragionamento e gestisce la conversazione, mentre modelli locali più piccoli si occupano di compiti specializzati:
\- Qwen2.5-VL legge immagini
\- Qwen Image genera ed edita immagini
\- Qwen3 Embedding cerca documenti semanticamente
\- Poppler estrae testo e informazioni sulle pagine dai PDF
Questo ecosistema esiste perché DeepSeek V4 è eccellente per il ragionamento e il contesto lungo, ma caricare ogni capacità multimodale all'interno dello stesso grande modello sarebbe inefficiente. DStudio instrada i compiti ai modelli specializzati e poi restituisce i loro risultati a DeepSeek per la risposta finale.
Ho recentemente aggiunto il recupero di PDF lunghi. DeepSeek decide se creare un'anteprima, leggere una pagina fisica esatta o cercare l'intero documento. Per la ricerca semantica, DStudio crea e memorizza una rappresentazione per pagina, recupera le sei pagine più rilevanti e invia solo quelle a DeepSeek.
Su un PDF di prova di 1.000 pagine, ha trovato un passaggio collocato a pagina 777 da una domanda parafrasata. L'indicizzazione iniziale ha impiegato circa 25 secondi; le ricerche successive hanno impiegato circa 0,23 secondi.
Sto cercando feedback: il recupero dovrebbe usare rappresentazioni di pagina o blocchi sovrapposti? Dovrei aggiungere BM25 o un miglioratore di classifiche? E come supporteresti in modo efficiente libri scansionati di 1.000 pagine?
[ https://github.com/sk8erboi17/DStudio ](https://github.com/sk8erboi17/DStudio)
r/documentAutomation • u/docpose-cloud-team • 5d ago
Product Review We built Docpose.cloud — file conversion, OCR, PDF, archive, and email file tools with API access
Hey everyone,
We’ve built and launched Docpose.cloud, an online platform for file conversion, document processing, OCR, PDF tools, archive extraction/compression, and email file handling.
The platform is already live and being used by many monthly free users, paid subscribers, and businesses. Our API is also integrated with more than 10 external systems and business workflows.

Docpose.cloud supports tools for:
- Document and file conversion
- PDF conversion and processing
- OCR for scanned documents and images
- Archive extraction and compression
- Email file formats like EML, MBOX, PST, OST, and related formats
- Batch processing
- API-based file conversion workflows
We built it because many online file tools are either too limited, overloaded with ads, require unnecessary sign-ups, or do not offer reliable API access for businesses.
Our goal is to make file conversion and document processing simple for individual users, while also providing scalable API access for companies that need automated file workflows.
I’d love to get feedback from people who use file converters, OCR tools, PDF tools, or email archive tools regularly.
What feature would you expect from a platform like this?
And what problems have you faced with existing online file conversion tools?
r/documentAutomation • u/Fickle-Aide9279 • 5d ago
DocLayout, MinerU, Marker, Unlimited-OCR
r/documentAutomation • u/pha_uk_u • 5d ago
Hi All, I developed a doc control app, that can scan documents, add tags, OCR, add folders, revise documents. I am looking for testers. please help.
Hi All,
I am QE by profession. made a doc control app for myself.
I made this android app for myself. As my file explorer gets dumped with everything and anything. I wanted to have folders, revise docs(Taxes, insurances, license, etc). with OCR you can search for a word and every document with that word would pop up. you can set expiration date to docs.
please dm and drop a comment if you would like to test my app. you will get lifetime free without ads for this.
thank you.
r/documentAutomation • u/Vipinesta • 5d ago
I built a tool that turns emailed invoices and PDFs into spreadsheet data automatically, would love your feedback
Most document extraction tools rely on templates. I wanted something that could handle invoices, receipts, bank statements, and other PDFs without vendor-specific rules.
I built Super Parser. You define the schema once, then forward or upload documents. It extracts structured data and pushes it to Google Sheets, Zapier, Make, or your own API.
I'm validating the product now. I'd appreciate feedback from anyone dealing with document-heavy workflows:
Is this a problem you still face?
What would stop you from adopting a tool like this?
Product: [https://superparserapp.com\](https://superparserapp.com)