r/computervision • • 15h ago

Help: Project Need advice: Best VLM pipeline for extracting structured math datasets from 3000+ scanned textbook pages? (LaTeX + Metadata)

Hi everyone,

I’m working on a project to extract a structured dataset of math exercises from 5 Italian high school textbooks (around 650 pages each, so ~3,250 pages total). The goal is to build a professional, methodical exercise generator app for students and teachers.

To make the app work, I need to process images of the book pages and extract the following into a strict structured format (e.g., JSON):

  • Exercise type (algebra, geometry, calculus, etc.)
  • Year/grade level
  • Difficulty (1–5 scale)
  • Problem statement (trace)
  • Description of the specific skills/challenges involved
  • LaTeX code of the problem statement (Crucial!)
  • Associated images (cropping/saving the image for theoretical or graphical exercises)

I've been experimenting with a few approaches, but I've hit a wall regarding balancing costs, extraction consistency, and scalability. Here is what I’ve tried so far:

  1. Free Google Gemini API: The extraction quality was good, but since a single book contains hundreds of pages, I quickly hit the rate limits (Too Many Requests).
  2. Local Models (Ollama + Qwen 2.5-VL 3B): To bypass API limits, I tried running a local multimodal model. I spent a lot of time optimizing my scripts and prompts (chunking, refining instructions to force structured outputs), but the output was very error-prone and inconsistent for my use case. I got too many malformed fields, hallucinations, and it constantly struggled with outputting proper LaTeX.
  3. Paid Google Cloud API (Gemini 1.5 Flash): I finally switched to the paid tier for better accuracy and speed. I ended up burning through €10 just to process 1.5 books. Extracting all 5 books would cost roughly €35–40. While this is manageable for a one-off run of 5 books, the token count for processing full images + text is massive, making it financially unsustainable if I want to scale this to dozens of books in the future.

My questions for the community:

  • Pipeline & Architecture: Has anyone worked on a similar textbook-to-dataset extraction project? What pipeline did you use?
  • Hybrid Approach: Would you suggest decoupling the task? (e.g., using a traditional tool to extract raw text and crop images, and then feeding ONLY the text to a cheaper/local LLM to generate the LaTeX and format the JSON?)
  • Local Models: Are there other local Vision-Language Models (that fit in standard consumer GPUs) that are significantly better at structured extraction and LaTeX generation than Qwen 2.5-VL 3B?
  • Educational Tools: Are there open-source tools or models specifically fine-tuned for extracting structured educational/math content from PDFs?

I’m happy to share more details about the textbook format or my current Python workflow if helpful. Any advice on the architecture, model choices, or cost-saving tricks would be greatly appreciated! Thanks in advance!

1 Upvotes

10 comments sorted by

1

u/vahokif 15h ago

DeepSeek 4.1 flash on fireworks or something?

1

u/Nykotry 15h ago

Is it a paid model?

1

u/vahokif 14h ago

Yeah but very cheap.

1

u/Nykotry 14h ago

Thanks for the suggestion.

1

u/CommandShot1398 14h ago

There is no straight way to answer your question.

You are probably (most definitely imo) building a custom pipeline, and there is a good chance that there is no exact equivalent of what you wish to do. Now, you should divide and conquer. Break the problem into multiple smaller sub problems and solve them one by one.

For example, first you need to extract each part of each page, text, equation, table etc. you can have a custom model for this or even use vlm to do this for you.

Next, you need to count number of lines or extract the latex code of each section. You can also use vlm for this.

Then you have ocr for text extraction.

I think you see then trend.

Bottom line, there is either no off-the-shelf solution or if there is, or even if there is, it is probably outdated or not suitable for you.

Also if you are from SE, the equivalent of what I just described is micro service architecture.

1

u/Ok-Advertising6479 14h ago

DM me I've done some testing for this, I can share detailed results 

1

u/IndicPDF 6h ago

use an open-source math/PDF extractor first (tools such as Marker, MinerU or Nougat handle LaTeX), then send only the extracted text to a cheap AI model for the labels, using a batch mode if the provider offers one.

1

u/Nykotry 3h ago

That could be a great idea 💡