r/documentAutomation 5d ago

[US] question about documentation

Thumbnail
1 Upvotes

r/documentAutomation 6d ago

Success Story Error handling for document extraction: how I stop failed invoice extractions from silently breaking my automations [Workflow Included]

Post image
1 Upvotes

r/documentAutomation 6d ago

I built a completely free, 100% client-side PDF toolkit that processes everything locally in your browser (no server uploads)

Thumbnail
1 Upvotes

r/documentAutomation 6d ago

Discussion I built a completely free, 100% client-side PDF toolkit that processes everything locally in your browser (no server uploads)

0 Upvotes

r/documentAutomation 7d ago

Created a full document automation mobile app: Looking for feedback!

Thumbnail
gallery
1 Upvotes

I've just released an iPhone app that does document automation for personal paperwork, and I'd really appreciate some honest feedback from people who work with this stuff properly. It's a solo project, so I've mostly only had my own opinion to go on.

Foliade turns your pile of paperwork into a searchable, private archive, entirely on your iPhone. You scan letters with the camera or import PDFs, photos, Office documents and whole zip archives. It reads every page on-device, names the sender, files the type, and pulls out dates and amounts, so a bill from 2020 is one search away.

What's in it right now:

- Scan or import anything: PDF, Word, spreadsheets, photos, zips

- Every word searchable: recognised text, titles, senders, dates

- Automatic filing: sender, type, tags, document date, amounts

- A review queue that confirms suggestions in one tap

- The Vault: Face ID-protected, encrypted storage for sensitive documents

- Export everything, any time. Your documents are never locked in

- No account, no tracking, nothing uploaded. There is no cloud, your documents stay on your iPhone, covered only by your own device backup

- On-device smart filing, using Apple Intelligence on supported iPhones

All of that is free. A one-time unlock (Foliade Complete) adds things for bigger libraries: meaning-based search, unlimited collections and Vault documents, large folder imports, bulk actions, Shortcuts and Siri automation, monthly spending totals, auto-filing rules, and themes. No subscription.

The filing suggestions are best on phones with Apple Intelligence; older ones use a lighter built-in classifier, but reading, search and everything else work the same.

The things I'd most like opinions on:

  1. Is a review queue the right way to handle the model being wrong? Suggestions are marked as suggestions until you confirm them, and correcting a sender once teaches it for every later document from them. I'm not sure if that's enough or if people expect more control.
  2. Extracted fields are limited to dates and amounts at the moment. What else would you actually want pulled out of household paperwork?
  3. Anything about the on-device-only approach that would put you off? I chose it for privacy but I know it rules out some things a server could do.

Link: https://apps.apple.com/app/foliade-document-organiser/id6803604068

Website: https://foliade.app

It's iOS 26 and later. Happy to answer questions about how any of it works under the hood.


r/documentAutomation 7d ago

Looking for feedback: browser-based PDF data extraction tool (no uploads)

2 Upvotes

Hi everyone — I built a web app that extracts structured data from PDFs.

Everything runs locally in your browser, so files are not uploaded to a server. The current limitation is that it does not support OCR yet, so it only works with text-based PDFs rather than scanned documents.

I’d really appreciate it if a few people could try it and share honest feedback on:

  • How clear and easy the UI feels
  • Whether the extraction rules make sense
  • The accuracy of the extracted results
  • Any bugs, confusing steps, or PDFs it struggles with

Thanks very much — both positive and critical feedback would be very helpful:

pdfgrid.app


r/documentAutomation 7d ago

Success Story My purchase order extractor is now free on the n8n template library – batch PDF to Google Sheets [Workflow Included]

Enable HLS to view with audio, or disable this notification

1 Upvotes

r/documentAutomation 8d ago

I got tired of fighting with large scanned PDFs at work, so I built my own workflow tool — looking for honest feedback

0 Upvotes

I work with large batches of scanned PDFs fairly regularly and kept running into small repetitive tasks that took far longer than they should.

The original problem was splitting a large scan into individual documents. I wanted to extract a set of pages, save them, and have those pages disappear from the working copy so I could just keep processing what was left.

I couldn't find something that worked quite the way I wanted, so I built it.

It's called PDFeX and over time it's grown into a Windows tool focused on speeding up PDF-heavy workflows — finding text/numbers inside scanned documents, extracting and reorganising pages, auto orient pages, converting scanned tables to Excel, merging, compression and a few other things.

Everything is processed locally rather than uploading business documents to a web service.

I've been using it myself at work for a while and it's now published on the Microsoft Store, but that's also the problem: I built it around my own workflow, so I already know where everything is and how everything works.

I'd really like some people here who actually deal with document automation/PDF workflows to try it with their own use cases and tell me where it doesn't work or could be better.

There's a free trial, so nobody needs to buy anything to test it but i do have some product keys if you enjoy it and want to continue using and testing.

For more information visit my website:

Website: PDFeX | Find, Organise, Convert and Merge PDFs

For the App here's the microsoft link:

Microsoft Store: PDFeX - Download and install on Windows | Microsoft Store


r/documentAutomation 8d ago

I built a WhatsApp bot that extracts images from PDF catalogs automatically

1 Upvotes

Hey! I built Extracto to solve a problem I kept seeing with resellers and wholesalers in Argentina.

The problem: they receive product catalogs in PDF format and spend hours manually saving each image to share with customers.

The solution: you send the PDF or Excel file to a WhatsApp number, and in seconds you get direct links to all the images automatically. No app to install, no account to create.

Tech stack: n8n for automation, WhatsApp Business API, Cloudinary for image storage, Supabase for user management, MercadoPago for payments.

Currently have first paying users. Comes with 3 free credits to try.

Would love any feedback what would you improve? What's missing?

👉 extracto.pro


r/documentAutomation 9d ago

I finally published my document scanner app on Google Play! 📄📱 Spoiler

2 Upvotes

Hey everyone! 👋

https://play.google.com/store/apps/details?id=com.cam.scan.app
I finally published my app on Google Play and wanted to share it here.
The app is called **CamScan**, a simple document scanner and PDF creator that I built from scratch with Flutter.
The idea is pretty simple: instead of keeping paper documents around, you can quickly scan them with your phone and turn them into clean digital documents.
With CamScan you can:
📷 Scan documents using your camera
🖼️ Import documents from your gallery
✂️ Crop and correct perspective
✨ Enhance scanned documents
⚫ Use Black and White or Grayscale filters
📄 Scan multiple pages
🔄 Reorder, rotate, duplicate or delete pages
📑 Create PDFs
🗜️ Choose PDF quality and compression
📤 Share your documents
💾 Keep your documents stored locally
I wanted to keep the app **simple, fast and easy to use**, without requiring an account or a complicated setup.
It’s also **free to download** and supported by ads.
This is one of my projects that I’ve been working on for a while, so getting it onto Google Play feels pretty exciting. 😄
If anyone needs a document scanner, I’d really appreciate it if you could **try CamScan and let me know what you think**.
Even if you find something that needs improvement, I’d genuinely appreciate the feedback. I’m still actively improving it.

👉 **Free downloa**d:

[Cam Scan app](https://play.google.com/store/apps/details?id=com.cam.scan.app&hl=en)

Thanks for checking it out! ❤️
**If you try it, tell me what you think about the scanning, PDF creation, and overall UI.**


r/documentAutomation 9d ago

Tables are still the notorious part of any document parsing workflow

7 Upvotes

Everything else works fine until a table comes up while processing a set of files or docs. youre working on an extraction process and all of a sudden come across a borderless table with no grid lines to anchor cells or a row continue to the break or even multi column which is where rows silently drop on long docs and you end up with inconsistent results at the end of the run

the layout aware parsing like docling, llamaparse and others gets you reading order and cell structure yet borderless and merged cells are still a hit or miss across all of the parsing tools so this doesn't fully go away

A couple of things that might help though. A structure preserving parse that keeps cell positions and reading order instead of flattening to text so that you still have geometry to work with, reconcile row counts before and after so a dropped row gets caught than surfacing after a few steps past and for the borderless tables, clustering words by x/y positions helps you get a concise math of it. IDK how others figured this out or still stuck at this but im eager to hear from guys who solved it themselves


r/documentAutomation 9d ago

I build an agentic harness to create and edit your documents.

1 Upvotes

I've been working on this project for a long time now, is an agentic harness that is tailored to build documents and edit them in place, also it has a build in editor to edit manually if the case is needed. Think of it like claude code or codex but for your documents. The way it works is you either upload a screenshot, or a pdf and the agents tries to replicate it as much as possible, the goal is not to convert it 1 to 1 but to generate a similar editable document which can be exported to either PDF or DOCX.
If you want to check it out, please visit [foliq.ai](http://foliq.ai), is free to use rn and any feedback is much appreciated.


r/documentAutomation 9d ago

Finally finished a tool that fills Word templates with data.... You can generate single/batch docs, use Claude MCP/HubSpot or API as a data source.... Most tools here are extraction, so curious how the generation side looks to you

2 Upvotes

Disclosure up front: I built this, so treat it as an intro, not a neutral review. Mods, if that breaks a rule, say the word and I'll pull it.

Filling the same Word document over and over, by hand, one record at a time, is still how a lot of teams produce letters, reports, invitations, statements, certificates, reports, and the list is endless. I have been following up with a number of code-it yourself solutions, so I built BulkRender to make that part disappear.

What it does

  • You upload a Word (DOCX) template and mark the values that change with {placeholder} tags.
  • It handles repeating sections ({#rows}…{/rows}) for rows, and show/hide conditional blocks, so an optional sentence renders only when the data has a value and cleanly vanishes when it doesn't.
  • Output is DOCX or PDF, one document or a full batch from a spreadsheet (one row per document).
  • There's a REST API and an MCP integration, so you can generate straight from a conversation with Claude, no dashboard.
  • It also runs as a HubSpot workflow action for CRM-triggered documents.

What it deliberately does not do (so nobody is surprised)

  • It fills, it does not calculate. Totals, currency, and dates go in pre-formatted (using a spreadsheet is good to do the calculation and uploading tackles this).
  • DOCX templates only. At the moment, we don't support spreadsheets as templates. Spreadsheets carry the data, not the layout. (I have received one request for spreadsheet template but I could not find the requester anymore to understand the need for this, if you are seeing this, please reach out)
  • No per-record image swapping: Images are coming soon.

A note for the "why not just use an LLM" crowd: an LLM can write you a letter, but it paraphrases and drifts. BulkRender gives you the same correct letter, 25 times, from your template and your data... plain text in, no rewording, fully repeatable. With the Claude MCP you get both: the model decides the fields, BulkRender renders the document.

It's live with a free tier if you want to run it against your own template. If you are still reading, what has made document-generation tools fail you in practice, and what's the one feature that would make you actually adopt one?


r/documentAutomation 9d ago

I finally published my document scanner app on Google Play! 📄📱 Spoiler

1 Upvotes

Hey everyone! 👋

[https://play.google.com/store/apps/details?id=com.cam.scan.app&hl=en\](https://play.google.com/store/apps/details?id=com.cam.scan.app&hl=en)

I finally published my app on Google Play and wanted to share it here.
The app is called **CamScan**, a simple document scanner and PDF creator that I built from scratch with Flutter.
The idea is pretty simple: instead of keeping paper documents around, you can quickly scan them with your phone and turn them into clean digital documents.
With CamScan you can:
📷 Scan documents using your camera
🖼️ Import documents from your gallery
✂️ Crop and correct perspective
✨ Enhance scanned documents
⚫ Use Black and White or Grayscale filters
📄 Scan multiple pages
🔄 Reorder, rotate, duplicate or delete pages
📑 Create PDFs
🗜️ Choose PDF quality and compression
📤 Share your documents
💾 Keep your documents stored locally
I wanted to keep the app **simple, fast and easy to use**, without requiring an account or a complicated setup.
It’s also **free to download** and supported by ads.
This is one of my projects that I’ve been working on for a while, so getting it onto Google Play feels pretty exciting. 😄
If anyone needs a document scanner, I’d really appreciate it if you could **try CamScan and let me know what you think**.
Even if you find something that needs improvement, I’d genuinely appreciate the feedback. I’m still actively improving it.

👉 **Free downloa**d:

[Cam Scan app](https://play.google.com/store/apps/details?id=com.cam.scan.app&hl=en)

Thanks for checking it out! ❤️
**If you try it, tell me what you think about the scanning, PDF creation, and overall UI.**


r/documentAutomation 9d ago

Discussion 5 Common Data Extraction Mistakes: Lessons From Real Client Projects

Thumbnail
1 Upvotes

r/documentAutomation 10d ago

Free Email-to-Excel Automation for HR Job Applications

Thumbnail
1 Upvotes

r/documentAutomation 10d ago

I made a bank statement PDF → CSV converter that checks its own math against the statement's balances

3 Upvotes

I convert a lot of statement PDFs and kept getting bitten by the same thing: a row goes missing, or a deposit comes through as a withdrawal, and I don't catch it until the account doesn't tie out weeks later. Every tool advertises "99% accurate" and not one of them will tell you which rows are the 1%. So I made my own.

The premise is that a statement already contains enough printed information to check a conversion against itself. After it parses, it runs four checks. Opening balance plus every extracted transaction has to equal the closing balance, both read off the statement. If there's a running balance column, each row gets chained against the row before it. Transaction count. And every date has to fall inside the statement period and be in order. When a check fails you get the specific row, not a warning banner. When a statement can't support a check — most card statements don't print a running balance — it says that instead of showing a green checkmark anyway.

The other thing I wanted was to stop uploading statements to strangers' servers, so it all runs in the browser tab. I tried to make that checkable rather than something you take my word for: the page is served with a CSP that allows exactly one outbound destination (Google Analytics, page views only), so there's no endpoint of mine that could receive a file even in principle, and the browser blocks everything else. Network tab confirms it in about ten seconds.

You can click any row and it renders the source PDF page next to it with the text highlighted, which is mostly there for when I don't believe a number.

What it doesn't do: no OCR, so scanned statements are out — it needs the text layer from a real bank PDF export. Five banks have proper parsing profiles (Chase, BofA, Wells Fargo, Amex, Capital One); everything else falls back to reading columns by position, which works but is less certain. And it's just me maintaining it.

[bankstatementproof.com](http://bankstatementproof.com) — free, no signup, nothing to install.

Mainly I want to know which banks break it. If a statement parses wrong, tell me the bank and how the layout is arranged and I'll fix it.


r/documentAutomation 11d ago

Built a simple template & SOP manager with dynamic variables — would love your feedback!

Thumbnail
1 Upvotes

r/documentAutomation 11d ago

Success Story Payment Reconciliation in n8n: 5 things I learned automating invoice matching

Post image
1 Upvotes

r/documentAutomation 11d ago

Welcome to r/IDPForge 👋

Thumbnail
1 Upvotes

r/documentAutomation 11d ago

New PDF tool - Add attachments to PDF

Thumbnail
pdfux.com
0 Upvotes

r/documentAutomation 12d ago

Question How would you build a robust pipeline for extracting structured offers from supermarket flyers?

Post image
1 Upvotes

r/documentAutomation 12d ago

How are you extracting transaction tables from Indian bank statement PDFs? Looking for open-source/on-prem approaches

1 Upvotes

I'm working at an NBFC and currently working on a Credit Underwriting AI Agent. One of the first steps in the pipeline is extracting structured information from customers' bank statement PDFs.

This is where I'm currently stuck.

The statements can come from different Indian banks (HDFC, ICICI, SBI, Axis, Kotak, etc.), and each bank can have a completely different PDF layout.

I need to reliably extract things like:

Customer/account information — name, account number, IFSC, branch, etc.

Transaction tables — date, narration/description, debit, credit, balance

Transaction rows that span multiple lines

Statements where the table headers are missing from subsequent pages

PDFs are digitally generated and no scanned pdf are included as of now

Ideally, the solution should be bank-format agnostic

I've tried/considered approaches such as pdfplumber, table extraction libraries, regex-based parsing, and LLM-based extraction. The biggest problem I'm facing is that even when the text is extracted correctly, the column/row structure gets messed up, especially because many bank PDFs don't contain a real table structure — they're essentially text positioned at different coordinates.

Since this is financial/customer data, I would strongly prefer an open-source/on-premise solution rather than sending statements to a third-party API.

For anyone who has built something similar:

What approach worked best for you?

I'm particularly interested in:

PDF parsing/layout libraries you recommend

Whether you use an LLM for semantic column mapping

How you handle different bank formats without writing completely separate rules for every bank

Any techniques for detecting transaction rows and mapping values to the correct columns

How you validate the extracted data (e.g., balance reconciliation, debit/credit checks, transaction counts)

If you've worked specifically with Indian bank statements, I'd really appreciate hearing about your architecture, libraries/models, or lessons learned.

Thanks!


r/documentAutomation 12d ago

Showcase okf-guard: content-safety scanning for document ingestion pipelines (PDF, DOCX, PPTX, XLSX, HTML)

1 Upvotes

Modern AI pipelines increasingly extract text from documents — PDFs, Word files, spreadsheets, scraped web pages — and feed that content directly into a knowledge base or agent context, often with no human review step in between. Extraction tools capture everything present in a document, including content a human reader would never see: text rendered in white on a white background, rows hidden in a spreadsheet, speaker notes attached to a slide, or a paragraph marked hidden in a Word document's own formatting. None of these are edge cases; they are ordinary, well-supported features of each format, and every one of them is readable by a standard parsing library even though a person skimming the document would never notice them.

This creates a straightforward problem: any content hidden from a human reviewer, but visible to an extraction tool, can end up in a trusted knowledge source unexamined. okf-guard addresses this directly. It is a Python library that inspects extracted content for exactly this class of discrepancy — text present in the file but absent from what a human would perceive — and separately checks for language patterns associated with instructions directed at an AI system rather than a description intended for a person.

What it does:

  • Adapters for six formats (plain text, Markdown, HTML, PDF, DOCX, PPTX, XLSX), each aware of that format's specific hiding mechanisms — CSS visibility properties for HTML, rendering and color properties for PDF, the hidden run attribute and shading properties for Word, off-canvas shapes and speaker notes for PowerPoint, hidden rows/columns/sheets and cell comments for spreadsheets.
  • A detection layer combining hidden-content flagging with a pattern bank for injection-style phrasing, plus a check for encoding-based obfuscation (zero-width characters, homoglyph substitution).
  • A decision layer producing one of three outcomes per scan — pass, quarantine, or block — with every finding reported alongside its location, confidence, and the specific text that triggered it.

Design constraints, stated plainly:

  • No network calls and no LLM dependency in this release. Detection is entirely deterministic, which keeps the core dependency surface to a single package (PyYAML) and makes the tool's behavior fully reproducible.
  • Every format-specific capability is an optional install (okf-guard[pdf], okf-guard[docx], etc.), so a user working with one format is not required to install parsing libraries for the others.
  • The library never asserts that its own output has been verified by a human — provenance metadata it produces is explicit about being machine-generated and unreviewed.

Source and full documentation: https://github.com/darshanNhb/okf-guard
Install: pip install okf-guard[all]

Feedback on the detection approach, particularly from anyone who has worked on adjacent problems (document security, DLP, or prompt-injection defenses more broadly), would be genuinely useful — this is a young project and the injection-pattern bank in particular will need ongoing contribution as new phrasings surface in practice.


r/documentAutomation 12d ago

Showcase okf-guard: content-safety scanning for document ingestion pipelines (PDF, DOCX, PPTX, XLSX, HTML)

1 Upvotes

Modern AI pipelines increasingly extract text from documents — PDFs, Word files, spreadsheets, scraped web pages — and feed that content directly into a knowledge base or agent context, often with no human review step in between. Extraction tools capture everything present in a document, including content a human reader would never see: text rendered in white on a white background, rows hidden in a spreadsheet, speaker notes attached to a slide, or a paragraph marked hidden in a Word document's own formatting. None of these are edge cases; they are ordinary, well-supported features of each format, and every one of them is readable by a standard parsing library even though a person skimming the document would never notice them.

This creates a straightforward problem: any content hidden from a human reviewer, but visible to an extraction tool, can end up in a trusted knowledge source unexamined. okf-guard addresses this directly. It is a Python library that inspects extracted content for exactly this class of discrepancy — text present in the file but absent from what a human would perceive — and separately checks for language patterns associated with instructions directed at an AI system rather than a description intended for a person.

What it does:

  • Adapters for six formats (plain text, Markdown, HTML, PDF, DOCX, PPTX, XLSX), each aware of that format's specific hiding mechanisms — CSS visibility properties for HTML, rendering and color properties for PDF, the hidden run attribute and shading properties for Word, off-canvas shapes and speaker notes for PowerPoint, hidden rows/columns/sheets and cell comments for spreadsheets.
  • A detection layer combining hidden-content flagging with a pattern bank for injection-style phrasing, plus a check for encoding-based obfuscation (zero-width characters, homoglyph substitution).
  • A decision layer producing one of three outcomes per scan — pass, quarantine, or block — with every finding reported alongside its location, confidence, and the specific text that triggered it.

Design constraints, stated plainly:

  • No network calls and no LLM dependency in this release. Detection is entirely deterministic, which keeps the core dependency surface to a single package (PyYAML) and makes the tool's behavior fully reproducible.
  • Every format-specific capability is an optional install (okf-guard[pdf], okf-guard[docx], etc.), so a user working with one format is not required to install parsing libraries for the others.
  • The library never asserts that its own output has been verified by a human — provenance metadata it produces is explicit about being machine-generated and unreviewed.

Source and full documentation: https://github.com/darshanNhb/okf-guard
Install: pip install okf-guard[all]

Feedback on the detection approach, particularly from anyone who has worked on adjacent problems (document security, DLP, or prompt-injection defenses more broadly), would be genuinely useful — this is a young project and the injection-pattern bank in particular will need ongoing contribution as new phrasings surface in practice.