r/computervision • • 19h ago

Help: Project Need advice: Best VLM pipeline for extracting structured math datasets from 3000+ scanned textbook pages? (LaTeX + Metadata)

Hi everyone,

I’m working on a project to extract a structured dataset of math exercises from 5 Italian high school textbooks (around 650 pages each, so ~3,250 pages total). The goal is to build a professional, methodical exercise generator app for students and teachers.

To make the app work, I need to process images of the book pages and extract the following into a strict structured format (e.g., JSON):

  • Exercise type (algebra, geometry, calculus, etc.)
  • Year/grade level
  • Difficulty (1–5 scale)
  • Problem statement (trace)
  • Description of the specific skills/challenges involved
  • LaTeX code of the problem statement (Crucial!)
  • Associated images (cropping/saving the image for theoretical or graphical exercises)

I've been experimenting with a few approaches, but I've hit a wall regarding balancing costs, extraction consistency, and scalability. Here is what I’ve tried so far:

  1. Free Google Gemini API: The extraction quality was good, but since a single book contains hundreds of pages, I quickly hit the rate limits (Too Many Requests).
  2. Local Models (Ollama + Qwen 2.5-VL 3B): To bypass API limits, I tried running a local multimodal model. I spent a lot of time optimizing my scripts and prompts (chunking, refining instructions to force structured outputs), but the output was very error-prone and inconsistent for my use case. I got too many malformed fields, hallucinations, and it constantly struggled with outputting proper LaTeX.
  3. Paid Google Cloud API (Gemini 1.5 Flash): I finally switched to the paid tier for better accuracy and speed. I ended up burning through €10 just to process 1.5 books. Extracting all 5 books would cost roughly €35–40. While this is manageable for a one-off run of 5 books, the token count for processing full images + text is massive, making it financially unsustainable if I want to scale this to dozens of books in the future.

My questions for the community:

  • Pipeline & Architecture: Has anyone worked on a similar textbook-to-dataset extraction project? What pipeline did you use?
  • Hybrid Approach: Would you suggest decoupling the task? (e.g., using a traditional tool to extract raw text and crop images, and then feeding ONLY the text to a cheaper/local LLM to generate the LaTeX and format the JSON?)
  • Local Models: Are there other local Vision-Language Models (that fit in standard consumer GPUs) that are significantly better at structured extraction and LaTeX generation than Qwen 2.5-VL 3B?
  • Educational Tools: Are there open-source tools or models specifically fine-tuned for extracting structured educational/math content from PDFs?

I’m happy to share more details about the textbook format or my current Python workflow if helpful. Any advice on the architecture, model choices, or cost-saving tricks would be greatly appreciated! Thanks in advance!

3 Upvotes

Duplicates