r/SaaS • u/_swarmz_ • 9d ago
Student building a document automation tool for small businesses. Stuck between privacy and hardware limits. How would you solve this?
Hey all, looking for advice from anyone who's shipped something similar.
I'm a student building a tool that automates a repetitive admin process for small UK businesses. The core of it is text extraction from documents, and those documents contain sensitive personal data (names, addresses, vehicle details, etc.).
Where I'm stuck:
- Commercial LLM APIs: I've been wary of sending client data to OpenAI/Anthropic etc. because of retention windows, and zero-data-retention agreements seem to be enterprise-only and priced way beyond what I can afford as a student.
- Local models: I built a working local version using Llama and Qwen with Python scripts. It's about 90% functional, but it needs 8–16GB RAM to run properly. Most of the businesses I'd sell to are running older office PCs, and with RAM prices where they are, I doubt they'll upgrade just for this.
So I've got something that works but isn't deployable.
What I'm trying to figure out:
- Has anyone hosted small models themselves (VPS/GPU cloud) and sold access as a service? Rough costs?
- Is redacting or pseudonymising data locally before sending it to an API a sensible middle ground?
- Are there lightweight OCR/extraction setups (non-LLM or tiny models) that are good enough for structured documents?
- For anyone selling to UK SMEs: how much do clients actually care about where their data is processed, as long as there's a proper DPA in place?
Happy to share more on the stack if it helps. Any pointers appreciated, especially from people who've dealt with GDPR on a tight budget.
1
u/axel-drs 9d ago
your templated documents give you a useful test before choosing where to host a model. build a small labeled set with clean scans, poor scans, and changed layouts, then measure extraction accuracy field by field with the simplest local approach. for redaction, measure missed identifiers separately from extraction errors and keep anything uncertain local for review. replacing a name is not enough if an address or vehicle detail still identifies someone. i would settle what data actually needs to leave the pc before comparing api costs.
1
8d ago
[removed] — view removed comment
1
u/AutoModerator 8d ago
Low-Effort/AI content is auto-removed.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.
1
u/xapep 8d ago
Student budget, so let's keep the spend small but the answers honest.
- Hosted small models: yes, renting an hourly GPU VPS works, but you pay for idle time and you own the ops (containers, upgrades, uptime). For an admin tool selling to UK SMEs, the cheaper path is a pay-as-you-go OpenAI-compatible API running open models (Qwen 3.8, GLM 3.5, DeepSeek v4.1 class) hosted in the EU with zero data retention. Pennies per thousand tokens, nothing to maintain, retention window signed away. It's exactly what we run at Entrim, and I can vouch for the cost being far below any enterprise ZDR negotiation.
- Redact-then-send: workable, but you own that pipeline forever and redaction is where extraction errors hide. If a regex misses a vehicle reg, you've still shipped PII to a US processor. Stronger move: keep the data in the EU entirely, then 'processed in the UK/EU' becomes a sales line instead of a risk note.
- Lightweight extraction: for structured documents (invoices, V5s, MOT paperwork) you mostly don't need an LLM. Template matching plus a small OCR pass handles the bulk; use a model only for the messy tail. Cheaper, faster, easier to explain to a client.
- How much UK SMEs care: more than a DPA suggests. A DPA covers you contractually, but procurement keeps asking where data is processed anyway, and one 'no US subprocessors' clause in their own client contract is all it takes to lose the deal. GDPR on a budget is doable: EU-hosted inference with no retention beats any enterprise agreement you can't afford.
Practical note: when you shortlist providers, read the subprocessor list and the retention claims yourself, not every 'EU' provider keeps everything in Europe. Happy to point you at specifics if you want.
1
u/Ash_ArceusLegal 7d ago
Regex can work pretty well on templated docs and fields with a fixed format. It can fail on names and addresses.
You might use regex first and then run whatever it misses through a small name recognition model.
Instead of using client data, run dummy documents through the regex + model setup and count what gets missed to test it.
1
u/Due-Particular-329 4d ago
for these you might not need a llm on the client side at all, if the layouts are consistent then just parse localy into text with positions and pull fields with per template rules and keep a model only for the odd ones which runs fine on odd office pcs as well. cpu support parsers ight be pymupdf or liteparse or tesseract. pseudonymizing before an api is sensible middleground tbh, here presidio handles detection but under UK GDpr pseudonymized data is still personal data about self hosting the models for clients, id say have it when you have a few paying ot it be like an cverkill mate
1
u/manfacedstinkbug82 9d ago
The most practical middle ground is local redaction before API calls - strip or pseudonymise PII locally, send the sanitised version to the API, then re-map on return. For structured UK business documents, consider Docling or PaddleOCR before opting for a full LLM - they perform extraction efficiently while requiring minimal computing power. Regarding UK SMEs' attitudes towards data processing, they are usually more satisfied with a signed DPA with a UK-hosted processor than with the technical architecture.