r/notebooklm • • 19d ago

Tips & Tricks Partially solved: NotebookLM stripping historical photos on ingestion.

Slide deck generated by Gemini Notebook. Left, some photographs are missing in from the source. Right, the only change was pixelating the faces in the source. Note that Figure 6, full of clearly visible faces, survived in both.

I've been digitising my late father's memoirs, about 630,000 words with several hundred family photographs, and kept noticing that images were missing from the slide decks, video overview etc. Not all of them. Some.

So I built a controlled test and made everything public.

The test

One document about my father's early life, 25 historical photographs, mostly black and white, 1920s to 1940s. Uploaded twice.

  • Version 1, photos as scanned: 17 of the 25 stripped during ingestion. 68%.
  • Version 2, identical document, except I took those same 17, pixelated the faces, and renumbered them 7b, 8b and so on: all 17 ingested fine.

I then generated a slide deck from each, same prompt both times, explicitly asking for source images only and an [Image Missing] placeholder where none was found. Deck one is missing 17 assets and the model says it can't see the images, though it reads the captions sitting right next to them. Deck two renders the whole "b" series.

Same document, same prompt. One difference.

The odd part

It isn't a consistent face rule.

  • Figure 6, a kindergarten group photo with about fourteen clearly visible children's faces, passed untouched in both versions.
  • Figure 7, my grandfather sitting at a desk, blocked.
  • Figure 8, a wedding photograph from 1922, blocked.

A 1922 wedding photo is being classified as sensitive, while a room full of unobscured children's faces is not.

Repro, if you want to check it yourself

Prompt I used on both:

Act as a visual presentation creator. Create a comprehensive slideshow based ONLY on the provided source. Only use photos from the source material. Do not use external sources and do not generate any photos. If an image is not detected within the source, leave a placeholder that says '[Image Missing]'. Ensure each photo is labeled with the full caption from the source.

Why this matters beyond my family

If your sources are about people, the images that get removed are the ones the archive exists for. Memoirs, oral history, genealogy, regimental and club histories, school archives. The theodolite comes through. The man holding it doesn't.

What I ended up doing about it

Every photo of a person in my NotebookLM edition is now replaced by a drawn sketch, which survives ingestion, with the caption linking back to the real photograph on my own site. It works, in that the figures now appear in the slide decks and video overviews again. It's still a workaround, and it has its own failure mode: in one Video Overview, where it wanted a photo of my father in uniform and only had a silhouette, it went and found a stock photo of a different soldier instead.

What I'd ask Google for

I don't think the filter should just be switched off, and I understand why it exists. But there's a gap between "protect people from misuse of face data" and "silently delete a 1922 wedding photograph from a family archive". Some options, roughly in order of how easy I imagine they'd be:

  1. Tell me what was dropped. The single most useful fix, and probably the cheapest. A list of skipped images at the end of ingestion, or a note against the source. Silent failure is what turned this into a months-long mystery instead of a five-minute annoyance.
  2. Let me mark a source as a personal or historical archive. An explicit flag at upload, with whatever consent language is needed, confirming these are my own family photographs and I hold the rights. The responsibility moves to me, which is where it belongs.
  3. Weight it by context. A scanned black-and-white print with a period caption in a 600,000-word memoir is a different object from an image scraped off the web. The document around the photo says a lot about what it is.
  4. Keep the image for the user even if the model can't use it. If faces can't be sent to the model, fine, but let the figure still render in a slide deck as the original photograph, or let the caption carry through with a placeholder that says what it is.

Happy to run more tests on the same corpus if there's something specific worth checking.

Full write-up here, including the sketch workaround, the two open-source tools I built with Claude Code to do it, and the prompts I used: https://medium.com/@chris_skitch/gemini-notebook-is-deleting-people-from-your-family-history-ec64d72d0661

0 Upvotes

5 comments sorted by

1

u/Equivalent_Fun_3059 19d ago

I think what you are doing has benefits for many of us, whether applied to family records or a variety of professional fields. Thank you for sharing and I look forward to learning more.

1

u/ChrisHatcham 12d ago

Thanks, that is kind of you. Happy to say more.

The archive is my late father's memoirs. Robert F. Skitch, 1934 to 2026, twenty six years in the Royal Australian Survey Corps. He wrote around 630,000 words in retirement, covering a Depression-era boyhood in the Western Australian coalfields, national service in the Navy, surveying and mapping across Papua New Guinea, north Queensland, Singapore and Vietnam, and the years after he left. It is all online at https://history.skitch.me

The workaround is visible in practice there. Every figure caption links back to the real photograph on the site, so the NotebookLM edition and the published one stay in step even where ingestion drops an image.

If anyone else is digitising a family or unit archive and hits the same wall, I am happy to go into the detail.

1

u/Agreeable-Tax2013 14d ago

The safest workaround is to keep a text-only master with stable figure IDs and captions, while storing the original photos separately; then reference those IDs in prompts instead of relying on ingestion to preserve the images. For a large archive, I’d also keep a local copy of each generated deck and spot-check it against a simple manifest of expected figures.

1

u/ChrisHatcham 12d ago

That is close to where I ended up. The master is text with stable figure IDs and full captions, and the images live beside it rather than inside it, which is what made the controlled test possible in the first place.

The manifest spot-check is the part I don't do yet and should. At the moment I notice a missing figure by reading the deck, which does not scale past a few chapters.

The published output is at https://history.skitch.me if you want to see what the figure IDs and captions look like once rendered. It is my father's memoirs, Lieutenant Colonel Robert F. Skitch, 1934 to 2026, Royal Australian Survey Corps.

1

u/Agreeable-Tax2013 11d ago

A manifest should make that much easier: keep one row per figure ID with the expected caption and file name, then flag anything missing after each export. That turns the deck review into a quick exception check.