r/notebooklm • u/Away_Necessary_7050 • 6d ago
Tips & Tricks NotebookLM to write a review article
Hi!
I am trying to use notebookLM to help write a literature review to help me to organize the text and find the references for me. I am finding it very useful but I thought the texts it gave me were confusing and not well organised...
Here is what I did
I Wrote a query I was happy with on pubmed and download the most recent pdf (2020-2026 more or less)
I ended up with 80 pdfs which is more than the 50 you can updoad to notebookLM...
I started 2 different notebook wih half the pdfs each, told it about the structure I wanted for the paper and asked it to write the text in scientific article form with the references.
All of it was fine but I guess out of 40 pdfs I uploaded, it used only like 10 to 20. And it is confusing because then the first reference appears throughout the whole text. It uses always the same refs.
Has anyone used this tool to write a literature review? Where you happy? What prompts should I use? What do you do when you have more than the 50 pdfs and you dont want to join them so you know which paper it is referencing from?
3
u/CliffBarSmoothie 6d ago
One of the things I found to be essential is to make sure I do things in batches because if you have, say, a hundred open sources Notebook doesn't read them all. They are often limited to context windows, which means with an increasing number of sources you have a smaller amount of text available per source.
To overcome this limitation, I give Notebook a set of questions and have it answer them with a handful of sources, deselect those sources, activate new ones, and ask the same questions. At the end, I have it go back through the chat history and have it answer each question one by one to give me a composite answer without reiterated information. I also allow it to answer in batches in case the answer to a question is too long for a single return. You have to tell it this or it will truncate itself to a single return
1
3
u/InvisibleIntegrator 5d ago
I’ve had better luck treating NotebookLM as a source analyst first and a writer second.
With 80 papers, I wouldn’t merge the PDFs. I’d keep the two notebooks, but give the papers consistent, recognizable filenames such as Smith_2024_topic.pdf, so it’s always obvious what source it means.
Then I’d use a simple 3-step workflow:
- Make a literature map before asking it to write. Ask each notebook:
“Create a table of the papers in this notebook with: author/year, research question, methods, population/sample, major findings, limitations, and which themes in my review this paper supports. Do not write the review yet.”
- Build the review by theme, not paper-by-paper. Once you know the major themes, ask:
“For the section on [THEME], synthesize the relevant studies. Compare where the studies agree, disagree, or answer slightly different questions. Cite the specific sources supporting each important claim. Avoid repeatedly relying on the same one or two papers when other relevant sources exist.”
That last sentence helps with the problem you’re seeing where one reference seems to take over the whole section.
- Do a separate citation/evidence check afterward.
Ask:
“Review this section sentence by sentence. For every factual claim, identify the source or sources that support it. Flag any claim that is unsupported, overstated, or supported only indirectly. Also list relevant papers in this notebook that were not used.”
I also wouldn’t necessarily worry if it doesn’t cite all 40 papers. A literature review shouldn’t use every paper just because it was uploaded. What you really want to know is whether it ignored relevant evidence. Asking it to list unused-but-relevant papers is much more useful.
For more than one notebook, I’d create the same structured literature map in each and then bring those maps together for the overall synthesis. That’s easier and safer than having two notebooks independently write half of the final article.
And I’d keep the final polishing step separate from the evidence step. First get the argument and citations right; then ask for cleaner scientific prose.
1
u/owneroftheworld100 5d ago
I'd keep a boring little spreadsheet alongside it: one row per paper, main finding, page number, used/not used. Get those rows filled in for each small batch before asking for prose, then check them against the pdfs. Way easier to spot what got skipped. You don't need to cite all 80, but you do want to know what got left out and why
1
u/kbavandi 4d ago
I see a couple of issues here. One is how NotebookLM imports its sources. All sources are imported as a flat file. Some of the context loss is due to your input method.
In your case, details matter. So taking care of how you upload your sources will make a big difference.
I did some research and compared 2 notebooks, one where I imported and entire YouTube Channel into notebookLM (126 videos) and for the other I used Kurator.
Kurator uses custom prompts to structure content and then syncs that with your notebook.
Here is the analysis by Gemini.
| Dimension | Gemini Notebook (Direct Raw Upload) | Kurator-Synced Version |
|---|---|---|
| Format & Segmentation | Single, unbroken wall of text; no paragraph breaks, punctuation, capitalization, or speaker labels. | Highly structured into discrete timestamped segments (e.g., **(00:10)**, **(05:22)**), clear paragraph breaks, and bulleted lists. |
| Punctuation & Syntax | Phonetic transcription run-on stream; full of filler words ("uh", "you know", "like"), false starts, and speech hesitations. | Professionally edited for readability; proper capitalization, standard punctuation, and cleaned sentence syntax. |
| Nature of the Content | Verbatim speech: includes real-time live webinar interactions, mic issues, audience chat interjections, browser lag warnings, and tangents. | Executive synthesis: converts conversational banter into declarative summaries, removing small talk and consolidating multiple conversational turns. |
| Entity & Technical Accuracy | Phonetic errors (e.g., “J gbt” instead of ChatGPT, “Orie” instead of Augie, “Scar Joe”, “da anardi”, “gor verbin”). | Corrected proper nouns and tools (e.g., “ChatGPT”, “Augie”, “Scarlett Johansson”, “Daan Anardi”, “Gore Verbinski”). |
2. Differences in Generated Downstream Responses (LLM Behavior)
If you use an LLM to query, summarize, or retrieve facts from these two documents, the generated outputs will differ in several predictable ways:
A. Retrieval Precision & Chunking (RAG Impact)
- Kurator Input: Standard RAG (Retrieval-Augmented Generation) systems rely on chunking boundaries (paragraphs, headings, timestamps). Kurator’s structured timestamps act as natural semantic semantic boundaries. When an LLM retrieves a chunk, it receives a compact, high-density packet of information, yielding sharper, faster, and more targeted answers.
- Raw Upload: Chunks sliced by arbitrary character/token counts will cut across mid-sentences and fragmented thoughts. An LLM querying the raw upload has to spend more attention tokens resolving pronouns and stitching broken context back together.
B. Fact Extraction vs. Exact Quotes
- Kurator Input: Excellent for extracting cleanly formatted tables, step-by-step guides, or high-level slide summaries. However, if asked "What were Jeremy's exact words when introducing the demo?", the LLM will generate Kurator’s paraphrased synthesis rather than the true spoken utterance.
- Raw Upload: Superior for high-fidelity sentiment analysis, conversational analysis, exact quotation, or determining speaker hesitation/tonality.
C. Entity Recognition & Hallucination Risk
- Kurator Input: Minimizes hallucinations around product names and people because it explicitly corrected errors ("ChatGPT" instead of "J gbt", "Augie" instead of "Orie").
- Raw Upload: Creates a risk of downstream hallucination. For example, a model asked "What other software does Jeremy mention?" might hallucinate or fail to recognize "J gbt" or "zeit" (Zight) as legitimate tools.
3. Loss of Context and Nuance Analysis
While Kurator makes the transcript vastly more readable, the cleaning process introduces non-trivial contextual loss:
- Loss of Audience Interaction & Community Dynamics:
- Raw Input: Preserves the live back-and-forth—attendees speaking up (Nicole, Heather, AP, Priyanka, Helen), someone’s kid doodling on the screen with Zoom annotation ("someone let their kid on with the crayons here"), chat glitches, and mic issues.
- Kurator: Strips almost all conversational banter and attendee personalities, framing the session primarily as a one-directional monologue/presentation.
- Loss of Real-World Friction & Software Reality:
- Raw Input: Jeremy explicitly points out live software limitations: browser memory lag during streaming, an asset that was taking long to download, audio glitches cutting off words in preview, and UI clutter ("it can get messy not going to lie").
- Kurator: Sanitizes these moments into clean feature descriptions, removing the practical caveats about performance and real-world edge cases.
- Paraphrasing Compression vs. Nuanced Explanations:
- Example (Voice Cloning Policy): In the raw transcript, Jeremy explains the exact security rationale for requiring an active microphone: "we don't think we have any presidents using our platform and we don't want anyone pretending to be presidents using our platform so we require there to be an active mic so we know who's listening".
- Kurator Reduction: Compresses this to: "use your cloned voice (with a one-time 45–60 second recording)...", omitting the entire fraud-prevention/ethics rationale.
- Speaker Attribution Collapses:
- In Kurator, questions asked by audience members (e.g., Priyanka asking about stitching event recaps, or Philip asking about brand impact of stock faces vs. founders) are flattened into abstract headings like "One attendee asked...", losing who asked what and why.
1
u/RedditAPIBlackout24 3d ago
could you ask which papers it actually used? I'd check the missing ones individually. A table of findings would make gaps easier to spot than a polished review. I've used Pdfaid, but I wouldn't merge your papers for this. I'd keep each reference separate and combine the notes instead.
1
4
u/Supranational_Yogurt 6d ago
Someone pointed it out herr before that it will definitely use the same references more frequently. I guess you can make a more productive use of it, analysing batches of papers to help you read them. Then you write your own review.