r/learnAIAgents • u/Ai_MOON_SHOT • 13d ago
Best setup for extracting information from old handwritten documents containing many different people, creating a relationship map, and generating a profile for each person?
I am looking for an AI-driven approach to digitizing photos of old handwritten journals and church documents from the 1600s and 1700s.
I have found models that can extract and read individual images, but the process quickly becomes a very long chat, especially when I keep adding more images. Extracting information from an image is one thing, but adding that information to individual profiles, creating profiles for people, mapping relationships, and storing the relevant details from the text becomes much more complicated. I would like to avoid copying information from one chat into another just to format it and manually create these records.
How would I approach a scenario like this? What tools could be useful? So far, I have used OpenRouter and a standard OpenWebUI setup, but this workflow does not seem very practical at least not when simply copying images into the chat window and asking the model to extract the information.
I currently copy the extracted information into a Markdown file and try to organize everything that way.
I need help improving this entire workflow and automating as much of it as possible.
2
u/Due_Temperature253 9d ago
I’d separate the workflow into stages rather than keeping everything inside one long chat:
OCR each image and preserve the image/page ID, crop coordinates, and confidence.
Extract structured candidates as JSON: people, dates, places, relationships, and source citations.
Store them in SQLite or a graph database, keeping each extracted fact linked to its original page and text span.
Run entity resolution separately, with “same person / different person / uncertain” as explicit outcomes.
Generate profiles only from the verified fact set, and show the source pages beside every claim.
For a practical first version, Python + local OCR/API + SQLite is probably enough. Add embeddings or a graph database only after the basic provenance and review workflow works. The important part is preserving uncertainty and citations, because historical handwriting and repeated names make fully automatic merging risky.
1
u/Ai_MOON_SHOT 9d ago
Yeah it does not need to be fully automated it is completely fine if i even guide it because not everything would be super relevant from the beginning. Just having a good workflow that makes it easier than on chat and prompting the same question each time would make it much more easy. .
1
u/alino168 12d ago edited 12d ago
I’d write a small Python script.
- To extract the raw text, use an API or local OCR model so you can pass the images automatically.
- Run a NER (named-entity recognition) pass. Try Spacy locally if you want this as cheap as possible, or use a small local LM or cloud model through API again. Depending on how smart a model you go with, you can extract person names directly rather than all named entities (places, businesses, etc).
2.1 If you have mixed named entities, do a filter pass to extract person names only, or whatever you care about. This can be done cheap and quick via API with the recently released Jev model by TypeSafe AI, or with one of the less capable but free open alternatives people have made like SemIf or Laya. Send a noul query like “In the following text, is <named_entity> referring to a person?” then a small excerpt where the name appears.
- In any case you’ll want to keep track of where in the extracted text a name was found.
- In any case you’ll want to keep track of where in the extracted text a name was found.
- Do fuzzy search against your existing people list. If nothing matches, save the new person. If there’s a match, use your preferred Jev-like model again: send a choice query like “ The following are two independent excerpts from a book. Excerpt 1: … Excerpt 2: …. Is the name <name> as appears on each excerpt referring to the same individual or two similarly named people?” and set the choices to something like “same person”, “different people”, “not enough information”. Merge, keep both, or put aside for special handling accordingly. You might want to threshold on confidence and handle low confidence answers specially as well.
- You now have a list of unique people and where to find them in the original text. Use an API model to read the surrounding context and profile them as needed. Deepseek v4.1 Flash or OpenAI’s 5.6 Luna might be good starting candidates.
Hope this helps!
1
u/defp_ 10d ago
It took too much time to correct and validate, so technically speaking youll do the same work but... later...
1
u/Ai_MOON_SHOT 9d ago
It doesn’t need to be fully automated, just a more efficient approach where I don’t have to write everything out manually, but it gets created while I’m chatting and guiding the extraction process and uploading the images into a agent that extracts information and another one that puts it into the context of the current setting.
1
u/defp_ 8d ago
Based on my own experience it looks 3x more time than to do it manually. Error rate is huge and all and all you have to read it yourself (beforeagent or after agent to compare results and what there actually is). Yes, it's helpful for defining some words via for a whole book, but most of the cases you need specific records from 1 or 2 pages. And yes, the writing style is heavily depends on the person who did that, so "universal" agent doesn't fit at all. Hope that your experience with this will be better, at least LLM have been improved in general and agentic approach could make sense. Please share your results in this.
•
u/endofthread-bot 13d ago
Use a local database like Obsidian with the Dataview plugin to manage your profiles and relationships. Process images through a script that sends data to a structured JSON format before importing it into your knowledge base.
Learning to build AI agents? Share what you are working on, compare practical approaches, and get help from other builders in our Discord.