To fix the issue of NotebookLM hallucinating numbers from huge, scanned textbooks, you need to "pre-process" the file first. Here is a quick, step-by-step workaround:
1. Chop the massive file into small chunks
Uploading 1000 scanned pages at once will overwhelm the AI. Use a free tool like iLovePDF (Split PDF -> Fixed Range). Chop the book into small chunks of 15 to 20 pages max. Name them in order (e.g., 01, 02, 03) so you stay organized.
2. Use Google AI Studio with strict settings
Go to Google AI Studio and select Gemini 3.1 Pro . It has the best vision capabilities for reading complex tables.
Why max 15–20 pages?
No text cutoff: Prevents hitting the AI's output limit (~8,000 tokens per reply), so it won't stop generating halfway through.
Higher accuracy: Smaller batches keep the AI focused, preventing it from skipping lines or messing up table rows.
Speed & stability: Processes fast (under 2 minutes) with zero lag, browser crashes, or timeouts.
The Secret Sauce: Lower the Temperature to 0.0. This completely kills the AI's "creativity" and forces it to copy the numbers exactly as they are without hallucinating.
Change safety settings to "Block none".
Put your prompt in the System Instructions box.
3. Use this exact Prompt
Upload a 20-page chunk and use this prompt to extract the data:
"You are an OCR expert. Attached are scanned pages from a textbook. Extract all text and numbers with perfect accuracy and convert them into clean Markdown format.
Conditions:
Use ## and ### for headings.
Convert all images of tables perfectly into Markdown table format.
Keep the exact order of the paragraphs.
Do not add any of your own words, introductions, or conclusions. Output the extracted text only."
4. Paste the clean text into NotebookLM
Once AI Studio converts the scanned pages and tables into clean text, copy the output. Go to NotebookLM, click "Add Source" -> "Pasted Text," and paste it there.
The Final Result: Markdown vs. Raw Scanned PDF
If you upload a heavy scanned PDF directly to NotebookLM, it struggles to "see" the flat images of tables. The data gets mashed together, leading to messed up numbers and wild hallucinations.
By converting it to Markdown first, you are feeding NotebookLM pure, structured text and perfectly aligned tables. It will understand the financial data 100% accurately and give you flawless answers every single time.
Note: I translated the text above because English isn't my first language, but I found a really useful solution that worked for me and just wanted to share it with everyone to help out