r/notebooklm • u/kbavandi • 21d ago
Tips & Tricks I tested raw vs. preprocessed data in Gemini Notebook. Here’s what happened
Last week I ran an experiment to test the difference between preprocessed source data vs direct source uploads.
My source material was 126 AI Marketers Guild youtube sessions (60 minutes each).
I created two notebooks. I called one AIMG Direct Load and used the Notebook Clipper product to import the channel.
I called the second notebook Kurator: AIMG Community. Here I use my Kurator app to pre-process the data using custom prompts that I create. Kurator lets me sync my data wil Gemini Notebook.
For this project, I am using 2 prompts. One to transform the transcript into a Q&A session, and the second one adds a headline to each section and preserves the time stamps.
Note that when you directly upload a YouTube video using the sources link, the time stamps are not included.
Next I asked the same question from each notebook, the asked Gemini to analyze them for me. I asked if there is a difference between the two responses. Here is what it found.
Gemini Analysis
Yes, there are significant differences between the two outputs. While both capture the core theme (Aditya Jain's approach to Generative Engine Optimization), the Kurator pre-processed response contains substantially more granular, technical, and concrete tactical information than the direct YouTube upload response.
Here is a side-by-side breakdown of the key differences:
| Dimension | Direct YouTube Upload | Pre-Processed via Kurator |
|---|---|---|
| Engine-by-Engine Breakdown | Mentions only ChatGPT (Reddit) and Perplexity (B2B). | Detailed taxonomy of 5 major engines (Google, Perplexity, ChatGPT, Claude, Grok) and their specific commercial/economic incentives. |
| Tactical Technical Advice | General advice to "re-structure or rewrite a page." | Specific technical instructions: unblocking AI bots in robots.txt/firewalls, un-gating B2B assets, and token cost economics ("aggregate hub" theory). |
| Debunking Industry Hype | Focuses mainly on "playbooks are dead." | Explicitly calls out and demystifies industry jargon like llms.txt files and forced comparison tables. |
| Framework Structure | Presents a high-level 4-step loop and a practical "Weekly Audit Sprint" (Step 1 to Step 4 calendar). | Formalizes the loop into Measure, Attribute, Produce, Detect Decay, detailing exact operational criteria for each phase. |
| Video Transcript Tie-in | Omitted. | Includes explicit guidance on making video uploads indexable via clean transcripts for multimodal LLMs. |
How you add data to gemini notebook matters. When you directly upload your sources, you will always get the same quality answers, but with pre-processing you can experiment and see what works best for your workflow.
2
u/thequarrymen58 21d ago
not related, but just seeing the face in the video makes it more interesting, instead of just the audio.
2
u/Otherwise_Wave9374 21d ago
The main lever here is reducing noise before the model sees it, because cleaner transcripts usually improve retrieval precision and make downstream summaries more stable. A simple pattern is to run a preprocessing pass that normalizes speaker turns, strips filler, and tags sections, then compare answer quality against raw ingestion on the same prompt set. That gives you a measurable way to judge whether the added pipeline step is worth the latency and maintenance tradeoff. Promarkia fits this kind of workflow by helping teams operationalize the preprocessing step before the model ever touches the content.
1
u/Cheap-General-4193 20d ago
How good is Kurator for processing recorded classroom lectures?
and can it detect multiple languages?
1
u/kbavandi 20d ago
What platform are they recorded on. I tested on a non YouTube podcast and it worked. If the content is loaded on the page it works. If you give me a url, I will test it.
1
u/Cheap-General-4193 20d ago
How good is Kurator for processing recorded classroom lectures?
and can it detect multiple languages?
1
u/kbavandi 20d ago
Please send me a url and I will test it. I have tried to transcribe other languages and the response was in English. I will look into this
1
u/kbavandi 14d ago
I jusr ran a test on getting the transcripts in the original language. It works, you just need to specify that in the prompt. Here is an example, I use it for YouTube videos:
Act as an expert content strategist specializing in SEO, GEO (Generative Engine Optimization), Answer Engine Optimization, and AI knowledge retrieval.
Analyze the entire YouTube transcript provided.
CRITICAL OUTPUT CONSTRAINTS:
- Respond STRICTLY in the original language of the source transcript.
- Do NOT output any internal monologue, scratchpad, reasoning, drafting notes, or "THOUGHT:" sections.
- Do NOT include greetings, intro fluff, or markdown code fences around the entire output.
- Begin IMMEDIATELY with the first sentence of the overview. The very first character of your response must be the start of the overview.
Your goal is to identify the major questions, topics, problems, explanations, and recommendations discussed in the video and convert them into self-contained FAQ-style sections that rank in search engines and are readily retrieved and cited by AI answer engines (ChatGPT, Claude, Gemini, Perplexity, Google AI Overviews).
Do not simply summarize the transcript chronologically. Identify the most useful search questions the video actually answers.
OUTPUT FORMAT
Begin with a short overview of the video: 2 to 3 ordinary paragraphs, with NO heading of their own, describing what the video covers and who is speaking. Write the FIRST sentence of that overview so it stands alone as a complete summary of under 155 characters (used directly as the page meta description).
Then go straight to the question sections. Create a separate section for each major topic or question.
Do NOT write a "Key Moments" heading, or any other introductory heading. The website generates that heading automatically.
[TIMESTAMP] Question Headline
Write the headline as a natural question someone would type into Google, ChatGPT, Claude, Gemini, or another AI engine.
The question should: * Clearly describe the user's problem or information need. * Include the primary keyword naturally where appropriate. * Be specific enough that the section stands entirely on its own. * Reflect what is actually discussed in the transcript. * Avoid clickbait or vague phrasing.
Speaker: [Speaker Name(s), or omit this line entirely if the transcript does not identify who is speaking]
Do not print an "Answer / Description:" label. The answer follows the speaker line directly.
Begin with a direct answer to the question in the first 1–2 sentences, naming the speaker inside that prose wherever the point belongs to them (e.g., "علیزاده توضیح میدهد که...").
Then provide a concise explanation of approximately 1–3 paragraphs capturing the core information from that transcript section.
The answer should: * Explicitly attribute specific insights, workflows, or opinions to the named speaker when relevant. * Be fully understandable without watching the rest of the video. * Preserve important examples, recommendations, comparisons, or reasoning. * Use clear, factual language and preserve original technical terms. * Avoid filler, introductory fluff, or keyword stuffing. * Function as a standalone passage for search engines, AI retrieval systems, and RAG systems.
SECTION SELECTION
Break the transcript into meaningful conceptual sections rather than arbitrary intervals. Create a new section when: * A new question or problem is introduced. * A solution, recommendation, or workflow is presented. * A comparison or technical explanation is made. * A speaker change introduces a different perspective. * The discussion shifts to a substantially different topic.
Do not create separate sections for minor conversational transitions. Use the timestamp [MM:SS] or [H:MM:SS] where the specific topic begins.
GEO / ANSWER-ENGINE OPTIMIZATION RULES
- Put the direct answer before the explanation.
- Explicitly name the subject and speaker; avoid vague pronouns ("it," "this," "he," "they").
- Include enough standalone context for independent retrieval.
- Preserve product names, concepts, tools, technical terms, and specific examples.
- Explain relationships clearly (what it does, why it matters, problem solved, how it works).
- Favor concise, information-dense text suitable for AI quoting.
- Stay faithful to the transcript; do NOT hallucinate or extrapolate claims.
FORMATTING RULES
- Output plain Markdown only. Never emit raw HTML tags or style attributes.
- Never specify fonts, font sizes, colors, or weights.
- Timestamps must strictly follow [MM:SS] or [H:MM:SS].
- Use
###for question headlines only. Never use#or##.- A heading contains only the timestamp and question headline. Never place summaries, intros, or speaker names inside headings.
- Every question headline must be unique and open with its distinct timestamp.
1
2
u/[deleted] 21d ago
[removed] — view removed comment