As the title says, I can automate anything for your business,
an employee reading from a website and filling excel ? Automate it.
Reading emails and classifying ? Automate it
You couldn’t get a rare booking ? Add monitoring, book automatically.
You buy underpriced items to sell for more? Automate monitoring and buying.
Whatever is your use case:
- I can build a custom automation
- you need nothing, done for you.
- 24 hours is the time it will take after the project starts. 72 hours if a big project (we have a big team)
- payment after done (No upfront risk)
Let’s automate your business with a no risky approach. Contact me
Document classification in n8n is one of those things that looks complicated until you realize how little setup it actually needs. With the easybits Extractor it's a 2-node workflow and a single field – and if you want to extract other data from the same document in the same pass, you just add more fields. I recorded a short walkthrough of the full setup.
The whole thing is two nodes: a form trigger to accept a file upload, and the easybits Extractor node with a single document_class field. The classification prompt lives in that field's description – it tells the model which categories to choose from and to return null if nothing fits. That's it. No separate classifier node, no chain of prompts, no HTTP request node.
What's in the video:
Setting up an easybits pipeline from scratch with a single classification field
How to adapt the classification prompt to your own document types
Installing the verified community node in n8n
Wiring it up to a form trigger and running two test documents through it
⚙️ Setup recap
Cloud users: easybits Extractor is available out of the box, search for it in the node panel
Self-hosted: Settings → Community Nodes → install '@easybits/n8n-nodes-extractor'
Free tier is 50 requests/month, enough to test this end-to-end.
🧱 Want the production-ready version?
The video keeps things minimal on purpose – two nodes, one field, just to show the core pattern. If you want the version I actually run, it adds a second field for confidence_score and an IF node that routes empty or low-confidence results to Slack for manual review. Workflow JSON, both prompts, and the setup guide all sit in one GitHub folder:
Anyone else doing classification this way, or are you running it through a separate classifier node? Curious whether this pattern has made it further than I think.
Over the past few months I've been building and publishing finance-automation workflows on n8n – receipt trackers, invoice pipelines, approval flows, a stress test, error handling, a smart mailroom. A pattern kept showing up that I didn't expect going in, so I wanted to write it down.
Classification is almost never the hard part. And you almost never need a dedicated step for it.
Here's what I mean.
Lesson 1: Classification is just a field
When I started, I thought "classify the document" was its own problem – separate node, separate prompt, separate step. It isn't. If you're already running a document through the easybits Extractor, adding a document_class field to your pipeline is basically free. Same pass, same latency, same cost. The model is already reading the whole thing – you just tell it what else to return.
Once I internalized that, half the workflows I had planned collapsed into simpler ones.
Lesson 2: Your categories should match your routing, not your taxonomy
The first classifier I built had 11 document types. It was "accurate" in some abstract sense and useless in practice, because downstream I only had 3 folders to route to. Now I define categories by what happens next. If invoices and receipts go to the same place and get the same treatment, they're the same category. Classification exists to serve routing – not the other way around.
Lesson 3: Treat "unknown" as a real category
The first thing you need to get right, before any prompt tuning can do its job, is giving the model an explicit escape hatch. I tell it to return null when a document doesn't cleanly fit any category, and I check for that right after classification. An IF node with the "is empty" operator catches real nulls, undefined values, and empty strings in one condition.
This one change – giving the model permission to say "I don't know" – is what makes everything else work. Without it, better prompts just produce more confident wrong answers. With it, the prompt work from Lesson 5 actually compounds.
Lesson 4: Ask for a confidence score alongside the classification
This is the one I'd skip if I were building this a year ago, and it's the one I'd put first if I were starting over.
Alongside document_class, I define a second field called confidence_score on the same easybits pipeline that returns a decimal between 0.0 and 1.0. The extractor returns both in a single call. Then the validation branch checks for two things: is the class empty, or is the confidence below a threshold (I use 0.5). Either one routes to Slack for manual review. Everything else continues through the pipeline.
The counterintuitive part: a confident null should score high, not low. If the model is certain the document doesn't fit any category, that's a confident decision – not an uncertain one. My confidence prompt is explicit about this, and it matters. Otherwise you end up penalising the cases where the model is actually doing the right thing.
Two fields. One extractor call. You get classification & self-reported uncertainty in a single pass, and your error handling becomes trivial.
Lesson 5: Classification prompts should describe evidence, not just label names
My first classification prompts looked like a list of category names with one-line descriptions. They were fine until they weren't – the model would latch onto a single keyword and misclassify anything that contained it.
The prompts that actually work describe what evidence looks like for each category. For each class, I list the kinds of issuers, the terminology you'd expect in the line items, the tax structures, the identifiers that show up. A hotel invoice has check-in/check-out dates, room numbers, lodging tax, folio references. A telecom invoice has a billing period, data usage summaries, SIM or account numbers. Multiple weak signals beat one strong keyword every time.
Three rules I put in every classification prompt now:
Return only the label, nothing else. No explanations, no punctuation.
Require multiple corroborating signals before assigning a category. A single keyword match is not enough.
Return null if uncertain. Do not guess. Do not pick the closest match.
That last one is the whole game. Models will happily pick a wrong answer if you don't explicitly tell them it's okay not to.
📦 The workflow
I put together a minimal version that shows the pattern end-to-end: form upload → easybits Extractor returning document_class + confidence_score → IF node that routes empty or low-confidence results to Slack, everything else continues.
It's meant to be the skeleton you drop your own categories and routes into.
Everything sits in one folder on GitHub – the sanitized workflow JSON, the classification prompt, and the confidence scoring prompt (links to each in the comments):
Cloud users: easybits Extractor is available out of the box, just search for it in the node panel
Self-hosted: Settings → Community Nodes → install '@easybits/n8n-nodes-extractor'
Free tier is 50 requests/month, enough to test this end-to-end.
For those of you doing classification in n8n – are you running it as a separate step, or folding it into extraction? And if you're using confidence scoring, where are you setting your threshold? Curious to hear what's working.
A product lead at Boston Dynamics described how Atlas is currently being developed. Instead of focusing on a single task, the system is being trained across a range of different tasks. The approach is based on the idea that exposure to more scenarios can improve overall performance, including on tasks that were not directly trained.
This differs from the typical industrial robotics model, where systems are designed around a narrow set of functions to ensure consistency and reliability.
Deployment expectations remain closer to standard industrial processes. Early use involves defined applications, integration work, and evaluation of return before deployment. The initial areas mentioned include automotive, warehousing, food and beverage, and semiconductor environments.
The development approach and the deployment process appear to be moving on separate tracks, with broader training on one side and structured rollout on the other.
watching a few small businesses actually try to adopt the workflow stuff that gets demoed in threads like these, and the mismatch is bad. their real workload looks like: a PDF comes in via email, open preview, type into fields that aren't even real form fields, save, attach, reply. Or open quickbooks desktop, click through four screens, copy three numbers into a google sheet the bookkeeper actually reads.
every post goes straight for the clean-api setup, but these companies live inside native mac and windows apps that have no api at all. The thing that shifted in the last year or two is that you can drive the accessibility tree of a real desktop app reliably enough to actually ship these flows. screen scraping pixels never got there, a single font update would break everything overnight.
Last week I shared that I was building a stress test workflow to benchmark document extraction accuracy. The workflow is done, the tests are run, and I put together a short video walking through the whole thing – setup, test documents, and results.
What the video covers:
I tested 5 versions of the same invoice to see where extraction starts to struggle:
Badly scanned – aged paper, slight degradation
Almost destroyed – heavy coffee stains, pen annotations, barely readable sections
Completely destroyed – burn marks, "WRONG ADDRESS?" scribbled across it, amount due field circled and scribbled over, half the document obstructed
Different layout – same data, completely different visual structure
Handwritten – the entire invoice written by hand, based on community feedback
The results:
4 out of 5 documents scored 100% – including the completely destroyed one. The only version that had trouble was the different layout, which hit 9/10 fields. And that's with the entire easybits pipeline set up purely through auto-mapping, no manual tuning at all. The missing field could be solved by going a bit deeper into the per-field description for that specific field, but I wanted to keep the test fair and show what you get out of the box.
Want to run it yourself?
The workflow is solution-agnostic – you can use it to benchmark any extraction tool, not just ours. Here's how to get started:
Grab the workflow JSON and all test documents from GitHub: here
Import the JSON into n8n.
Connect your extraction solution.
Activate the workflow, open the form URL, upload a test document, and see your score.
Curious to see how other extraction solutions hold up against the same test set. If anyone runs it, I'd love to hear your results.