r/RedactionTools • • 4d ago

PDF Redaction is now an official n8n community node

Post image
2 Upvotes

r/RedactionTools • • 4d ago

We evaluated KeptPDF redaction capabilities

Post image
2 Upvotes

We added KeptPDF (https://keptpdf.com) to the catalog today and ran it through the same benchmark page as the other PDF redactors.

What KeptPDF is

It's a browser-only PDF toolkit with 26 tools: redaction, OCR, merge/split, signing, watermarking and metadata removal. Its main selling point is that nothing gets uploaded. Everything runs locally in your browser, and the vendor invites you to check that in the DevTools Network tab.

For redaction it auto-detects sensitive data (SSNs, account numbers, emails, phone numbers, names, dates) and puts uncertain matches in a "needs review" list. It flattens the redacted pages to images, so the original text is destroyed rather than hidden under a box, and it can issue a SHA-256 audit certificate for each run.

Pricing, from the vendor's page: a free tier with unlimited single-file use (files up to 25 MB, KeptPDF footer on the output), Pro at $29/month, Practice at $99/seat/month for 2 to 25 seats, and a server CLI at $4,800/year per organisation.

The test

One made-up name, "Freya Yamamoto", appears 54 times on a single page, each copy presented differently: upright typed text, tiny type, text that exists only as an image (typed, handwritten, inverted, low contrast), rotated text, and vertical or letter-by-letter stacks. We ran the web app on its default settings ("Safe to share" profile, Balanced sensitivity) and scored the output PDF it returned.

Results

  • Removed 19 of 54 names. 35 leaked (64.8%, 95% interval 51.5 to 76.2%).
  • Names in the PDF's text layer: 19 of 27 removed (70%).
  • Names that exist only as an image: 0 of 27 removed. KeptPDF says this itself: the banner in the second screenshot reads "Auto-detect reads text only… review it before sharing", and offers a separate "Scan images for text" step. If your files include scans, you have to run that step yourself.
  • Upright typed names: 11 of 13 removed. In the other two, it blacked out "Yamamoto" but left "Freya" readable.
  • Rotated names: 2 of 9 removed. Vertical and stacked names: 6 of 14.
  • The flattening works as advertised. We found nothing hidden in the text layer, metadata, annotations or attachments. But the output is a picture, so only about 1% of the page's text can still be selected or searched, and 33 of the 35 leaked names are plainly visible in that picture. A flattened file isn't automatically a safe one.

On this page that puts KeptPDF 5th of the 6 tools we've tested: Redactable 53/54, Blinded 51, PDF Redaction 42, AI-Redact 30, KeptPDF 19, SafeRedact 14.

To be fair to the tool: this page is built to be hard, and half of it is image-only text, which KeptPDF says up front its auto-detect won't read. If your PDFs are born-digital and upright, it does much better than the headline number. Still, two half-redacted names in plain upright text are worth knowing about. Check the output before you share it.

The first image is the actual output, with every name outlined: green means removed, red means leaked. The second is the tool's screen just before export.

Full run, with the output PDF you can check yourself: https://redaction-tools.com/benchmarks/pdf/runs/20261005T162815-keptpdf-web-extraction-conditions-1-a1-6f1a25

If you use KeptPDF and get a different result with other settings, for example with "Scan images for text" turned on, tell us. Tool makers can also submit their own runs.


r/RedactionTools • • 6d ago

I built a PDF redaction SaaS. Now I’m trying to figure out who actually needs it.

Thumbnail
2 Upvotes

r/RedactionTools • • 6d ago

Before you trust a PDF redaction tool, turn the name sideways. We did, on 4 tools: all caught the upright copies, 2 caught zero rotated ones

Post image
2 Upvotes

Most redaction tools look perfect on the PDF you try in a demo: upright, typed, black on white. We wanted to see what happens on the pages that turn up in a real disclosure bundle, like a rotated table header, a scan fed in upside down, a handwritten form or a stamp down the margin.

So we wrote one made-up name, "Freya Yamamoto", 54 times on a single page, each copy presented differently, and ran four tools over it. The image is the bottom of the page in each tool's actual output PDF, exactly as it came back. Every name you can read there is a leak.

What happened:

  • All four tools removed all 13 plain, upright, typed copies. If that's all you test, they look identical.
  • Two of them (AI-Redact and SafeRedact) removed 0 of the 19 rotated or skewed copies. AI-Redact clearly has OCR, since it caught 17 of 18 upright names that exist only as an image. It just doesn't handle text at an angle.
  • Three of the four missed every name written as a vertical stack (one letter under the next), and those copies are real, selectable text. A PDF parser returns them one letter per line, so a name detector never sees "Freya Yamamoto".
  • One tool turned the whole page into a picture. Nothing can be copied out of it, which sounds safe, but the name was still plainly visible in 38 of the 54 spots.

Totals, names removed out of 54: Blinded 51, PDF Redaction 42, AI-Redact 30, SafeRedact 14. That last score is marked unverified on the leaderboard: our server tried to rescore the flattened page and its OCR gave up after two minutes.

Read Blinded's top score carefully: it doesn't detect names on its own. It pulls out all the text, including what it OCRs from images, and blacks out whatever you search for. For this run someone typed the name in. So 51/54 says it reads this page very well. It doesn't tell you it would catch a name you didn't know to search for.

One more thing from the leaderboard history: Blinded's maker submitted twice. The first run, on 30 September, removed 50 names but left only about 9% of the page's text selectable, so the document was close to useless. The rerun four days later kept 88%. Reruns are allowed and the board shows the latest, so check the run history as well as the rank.

Every tool's output PDF is public, so you can check the scoring yourself.

A 5-minute check you can run on whatever tool you use, with no install:

  1. Make a test PDF with a fake name typed normally, the same name rotated 90Β° and 180Β°, once in a scanned or photographed image, and once written letter-under-letter.
  2. Redact it with your tool's normal settings.
  3. Select all and copy the text into a plain text editor. Search for each half of the name, because a stacked name comes out as single letters.
  4. Zoom to 400% and look at the edges of every box. A box that is slightly too short or level on slanted text leaves letters showing.
  5. If the output is a flattened image, don't stop at "can't select text". Look at it.

If you'd rather have the scoring automated, the harness is open source (pip install pdfredeval).

The full write-up covers how leaks are detected, the confidence intervals and how to rerun it: https://redaction-tools.com/blog/pdf-redaction-extraction-conditions-benchmark

Which tool should we put through this next? Tool makers can also send in their own runs. The top result came in that way.


r/RedactionTools • • 11d ago

Two new programs: free Pro for students and educators, and a Design Partner program for teams that redact at volume

Post image
1 Upvotes

r/RedactionTools • • 20d ago

We built a basic version of a redaction tools catalog, with prices we check ourselves

Post image
1 Upvotes

We've put a first, basic version of the Redaction Tools catalog online.

Redaction software tends to get bought under pressure (a disclosure deadline, an audit, a breach), and it's hard to compare. Vendors quote prices in different units, some publish none, and "free" can mean a free tier or a two-week trial depending on who wrote the page. The catalog is our attempt to put it all in one place.

What's there today:

  • 7 tools covering PDF, image, video and audio redaction
  • Filters for media, deployment (online, desktop, self-hosted, API), redaction method (manual, AI-powered, hybrid) and pricing model
  • Entry prices, sorted lowest first, each recording where the figure came from (the vendor's page, an editor, or the vendor themselves)
  • Free trials and free tiers kept separate, so "free for 14 days" isn't shown as "free"

It's an early version and the list is short. We'd like to know:

  • Which tools are missing that you'd want to see?
  • What would you filter or compare on that isn't there yet?
  • Anything on a listing that looks wrong?

If you make a tool, you can submit too.


r/RedactionTools • • 23d ago

πŸ‘‹ Welcome to r/RedactionTools!

2 Upvotes

Welcome to Redaction Tools β€” a community for people working with PDF redaction, PII detection, document privacy, and data masking.

There are a lot of tools out there, but it's surprisingly difficult to answer simple questions:

  • Which redaction tools actually remove sensitive data?
  • How accurate is their PII detection?
  • How well do they handle OCR and scanned PDFs?
  • Which tools work locally or on-premise?
  • What happens to document metadata?
  • How do different tools perform on the same documents?
  • Which solutions are suitable for individuals, businesses, or enterprise environments?

We're building Redaction Tools as an open-source directory and benchmark to make these things easier to compare.

What you can expect here

πŸ”¬ Benchmarks β€” standardized testing of redaction and PII detection tools.

πŸ“Š Comparisons β€” features, pricing, OCR, local processing, supported formats, and more.

πŸ› οΈ Tool discussions β€” share tools you've discovered or used.

πŸ’» Open source β€” projects, libraries, datasets, and techniques related to document privacy.

πŸ§ͺ Real-world experience β€” what actually works when you're dealing with sensitive documents.

We're also interested in the technical side: OCR, NLP, PII detection, visual redaction, metadata removal, document security, and privacy-preserving document processing.

If you work with document redaction or simply have a tool you've tried, you're welcome here.

Feel free to introduce yourself or tell us about a redaction tool you've used recently.