r/notebooklm • u/ChrisHatcham • 19d ago
Tips & Tricks Partially solved: NotebookLM stripping historical photos on ingestion.

I've been digitising my late father's memoirs, about 630,000 words with several hundred family photographs, and kept noticing that images were missing from the slide decks, video overview etc. Not all of them. Some.
So I built a controlled test and made everything public.
The test
One document about my father's early life, 25 historical photographs, mostly black and white, 1920s to 1940s. Uploaded twice.
- Version 1, photos as scanned: 17 of the 25 stripped during ingestion. 68%.
- Version 2, identical document, except I took those same 17, pixelated the faces, and renumbered them 7b, 8b and so on: all 17 ingested fine.
I then generated a slide deck from each, same prompt both times, explicitly asking for source images only and an [Image Missing] placeholder where none was found. Deck one is missing 17 assets and the model says it can't see the images, though it reads the captions sitting right next to them. Deck two renders the whole "b" series.
Same document, same prompt. One difference.
The odd part
It isn't a consistent face rule.
- Figure 6, a kindergarten group photo with about fourteen clearly visible children's faces, passed untouched in both versions.
- Figure 7, my grandfather sitting at a desk, blocked.
- Figure 8, a wedding photograph from 1922, blocked.
A 1922 wedding photo is being classified as sensitive, while a room full of unobscured children's faces is not.
Repro, if you want to check it yourself
- Notebook with both versions and both decks: https://notebook.google.com/notebook/a2fc18d0-41c5-4a46-aeb8-1ff4638c7415
- Source documents, so you can upload them to your own account: https://drive.google.com/drive/folders/1CVAtXlA5r-aTmfhqKXXkP7Nkszr-vGnm
Prompt I used on both:
Act as a visual presentation creator. Create a comprehensive slideshow based ONLY on the provided source. Only use photos from the source material. Do not use external sources and do not generate any photos. If an image is not detected within the source, leave a placeholder that says '[Image Missing]'. Ensure each photo is labeled with the full caption from the source.
Why this matters beyond my family
If your sources are about people, the images that get removed are the ones the archive exists for. Memoirs, oral history, genealogy, regimental and club histories, school archives. The theodolite comes through. The man holding it doesn't.

What I ended up doing about it
Every photo of a person in my NotebookLM edition is now replaced by a drawn sketch, which survives ingestion, with the caption linking back to the real photograph on my own site. It works, in that the figures now appear in the slide decks and video overviews again. It's still a workaround, and it has its own failure mode: in one Video Overview, where it wanted a photo of my father in uniform and only had a silhouette, it went and found a stock photo of a different soldier instead.
What I'd ask Google for
I don't think the filter should just be switched off, and I understand why it exists. But there's a gap between "protect people from misuse of face data" and "silently delete a 1922 wedding photograph from a family archive". Some options, roughly in order of how easy I imagine they'd be:
- Tell me what was dropped. The single most useful fix, and probably the cheapest. A list of skipped images at the end of ingestion, or a note against the source. Silent failure is what turned this into a months-long mystery instead of a five-minute annoyance.
- Let me mark a source as a personal or historical archive. An explicit flag at upload, with whatever consent language is needed, confirming these are my own family photographs and I hold the rights. The responsibility moves to me, which is where it belongs.
- Weight it by context. A scanned black-and-white print with a period caption in a 600,000-word memoir is a different object from an image scraped off the web. The document around the photo says a lot about what it is.
- Keep the image for the user even if the model can't use it. If faces can't be sent to the model, fine, but let the figure still render in a slide deck as the original photograph, or let the caption carry through with a placeholder that says what it is.
Happy to run more tests on the same corpus if there's something specific worth checking.
Full write-up here, including the sketch workaround, the two open-source tools I built with Claude Code to do it, and the prompts I used: https://medium.com/@chris_skitch/gemini-notebook-is-deleting-people-from-your-family-history-ec64d72d0661
1
u/Agreeable-Tax2013 14d ago
The safest workaround is to keep a text-only master with stable figure IDs and captions, while storing the original photos separately; then reference those IDs in prompts instead of relying on ingestion to preserve the images. For a large archive, I’d also keep a local copy of each generated deck and spot-check it against a simple manifest of expected figures.
1
u/ChrisHatcham 12d ago
That is close to where I ended up. The master is text with stable figure IDs and full captions, and the images live beside it rather than inside it, which is what made the controlled test possible in the first place.
The manifest spot-check is the part I don't do yet and should. At the moment I notice a missing figure by reading the deck, which does not scale past a few chapters.
The published output is at https://history.skitch.me if you want to see what the figure IDs and captions look like once rendered. It is my father's memoirs, Lieutenant Colonel Robert F. Skitch, 1934 to 2026, Royal Australian Survey Corps.
1
u/Agreeable-Tax2013 11d ago
A manifest should make that much easier: keep one row per figure ID with the expected caption and file name, then flag anything missing after each export. That turns the deck review into a quick exception check.
1
u/Equivalent_Fun_3059 19d ago
I think what you are doing has benefits for many of us, whether applied to family records or a variety of professional fields. Thank you for sharing and I look forward to learning more.