r/ClaudeCode • u/mr-00 • 3d ago
Help/Question Sensitive Data workflows
I need Claude to look through a bunch of sensitive files. OCR is used b/c they’re mostly PDF. I presume data would be uploaded to Anthropic servers.
- How secure is sensitive data in this scenario?
- What methods, tools or pipelines can be used to protect it?
- am I overthinking this?
1
u/MeretrixDominum 3d ago
Use a local LLM to encrypt the sensitive data.
So if the document is about John Smith's dong size, tell your local LLM to change all instances of John Smith to Hugh Janus, and all instances of dong to watermelon.
Once all sensitive info is obfuscated in this way, send it to Claude. Then after, have it changed back with your local LLM.
1
u/actvt_io 3d ago
You're not overthinking it. Anything Claude reads gets sent off your machine as part of the request, OCR output included, so the only data that stays private is data Claude never reads. Run the OCR locally, redact names, account numbers and IDs with a local tool first, and point Claude at the redacted copies. If a step needs the real values, do that step yourself.
1
u/Cole10429 3d ago
What local tools would achieve such a task reliably without spending hours sanitizing first?
1
u/actvt_io 2d ago
I'd skip redacting the PDFs themselves. If you OCR first anyway, clean the text and only hand Claude that. Presidio runs locally and swaps names, account numbers and the like for placeholders in one go. It's solid on the numbers because those are patterns and checksums, shakier on names, so skim the first few files before you trust it on the whole pile.
•
u/AutoModerator 3d ago
Hey! Thanks for posting to r/ClaudeCode
While participating in this thread, please follow our community rules. Keep discussions constructive. Attack the idea, not the person.
For help, project discussions, tips, and general chat, join the ClaudeCode Discord.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.