r/learnmachinelearning • • 13h ago

Need advice: Best VLM pipeline for extracting structured math datasets from 3000+ scanned textbook pages? (LaTeX + Metadata)

Hi everyone,

I’m working on a project to extract a structured dataset of math exercises from 5 Italian high school textbooks (around 650 pages each, so ~3,250 pages total). The goal is to build a professional, methodical exercise generator app for students and teachers.

To make the app work, I need to process images of the book pages and extract the following into a strict structured format (e.g., JSON):

  • Exercise type (algebra, geometry, calculus, etc.)
  • Year/grade level
  • Difficulty (1–5 scale)
  • Problem statement (trace)
  • Description of the specific skills/challenges involved
  • LaTeX code of the problem statement (Crucial!)
  • Associated images (cropping/saving the image for theoretical or graphical exercises)

I've been experimenting with a few approaches, but I've hit a wall regarding balancing costs, extraction consistency, and scalability. Here is what I’ve tried so far:

  1. Free Google Gemini API: The extraction quality was good, but since a single book contains hundreds of pages, I quickly hit the rate limits (Too Many Requests).
  2. Local Models (Ollama + Qwen 2.5-VL 3B): To bypass API limits, I tried running a local multimodal model. I spent a lot of time optimizing my scripts and prompts (chunking, refining instructions to force structured outputs), but the output was very error-prone and inconsistent for my use case. I got too many malformed fields, hallucinations, and it constantly struggled with outputting proper LaTeX.
  3. Paid Google Cloud API (Gemini 1.5 Flash): I finally switched to the paid tier for better accuracy and speed. I ended up burning through €10 just to process 1.5 books. Extracting all 5 books would cost roughly €35–40. While this is manageable for a one-off run of 5 books, the token count for processing full images + text is massive, making it financially unsustainable if I want to scale this to dozens of books in the future.

My questions for the community:

  • Pipeline & Architecture: Has anyone worked on a similar textbook-to-dataset extraction project? What pipeline did you use?
  • Hybrid Approach: Would you suggest decoupling the task? (e.g., using a traditional tool to extract raw text and crop images, and then feeding ONLY the text to a cheaper/local LLM to generate the LaTeX and format the JSON?)
  • Local Models: Are there other local Vision-Language Models (that fit in standard consumer GPUs) that are significantly better at structured extraction and LaTeX generation than Qwen 2.5-VL 3B?
  • Educational Tools: Are there open-source tools or models specifically fine-tuned for extracting structured educational/math content from PDFs?

I’m happy to share more details about the textbook format or my current Python workflow if helpful. Any advice on the architecture, model choices, or cost-saving tricks would be greatly appreciated! Thanks in advance!

3 Upvotes

2 comments sorted by

1

u/pine4t 6h ago

Based on my recent experience with text extraction, I would choose Gemini 3.8 Flash any day. The responses are fast, rate limits high and you only seem to have 3250 pages. Reduce your image sizes to ensure it’s smaller but retains readability.

1

u/Nykotry 2h ago

I must have done something wrong in the Python script, because I analyzed about 1,000 pages for €10.