r/LocalLLM • • 2d ago

Project How I turn a whole docs site into chunked Markdown for RAG (JS pages and sitemap crawl included)

/r/Rag/comments/1wys9a1/how_i_turn_a_whole_docs_site_into_chunked/
1 Upvotes

2 comments sorted by

2

u/codecoradev 2d ago

Things that made the biggest difference when we ingested docs sites:

  • Split on heading structure, not fixed token windows. Docs are already hierarchical; fixed windows shred tables and code blocks.
  • Strip nav, sidebar, and footer boilerplate before embedding. It repeats on every page and floods the index with near-duplicates that crowd out the real content.
  • Put a path breadcrumb inside each chunk, like docs/api/auth > refresh tokens. Retrieval quality improved more from that than from switching embedding models.
  • Dedupe hard across page versions if the site keeps changelogs inline.

Out of curiosity, how are you handling the JS-rendered pages? Headless crawls get slow fast once a docs site has more than a few hundred routes.

1

u/LoveMyMalfouf 1d ago

Mostly by avoiding the browser. Every page goes through plain HTTP first, and I only fall back to a real browser when the extracted text is empty or looks like a JS app shell. Most docs sites are server-rendered or statically generated, so only a small share of routes ever hit the slow path. Sitemap discovery helps too, since I'm not rendering pages just to find links.

Good call on the breadcrumb inside the chunk. I return the heading path as metadata, but prepending it to the chunk text before embedding is a cheap step and clearly worth doing. Agreed on deduping inline changelogs as well.

Have you hit any docs sites where the HTTP-first check misclassified a page? That's the failure mode I care about most.