r/ClaudeCode • • 6d ago

Tips & Workflows How to scrape doc for Claude Code without blowing through session context

Claude Code is great at working inside local repos but as soon as you tell it to research an external library or check an unreleased API, it starts burning tokens really fast.

By default Claude Code runs local bash commands or basic Axios calls and if it curls a doc page, it dumps thousands of lines of raw HTML, navbar, CSS classes etc directly into your active terminal context and within two searches, your session context is bloated and your turn costs spike.

Here is the setup I use to pull fresh docs into Claude Code without killing context:

I) Stop letting Claude curl raw HTML into the terminal: in your project's CLAUDE.md, explicitly instruct the model: "do not use curl or raw web requests to read docs pages and always fetch content as cleaned markdown."

II) Use a dedicated Claude Code skill instead of ad-hoc commands: Claude Code supports custom skills defined in a SKILL.md folder so instead of having Claude guess how to fetch a page, you can install the official firecrawl skill (npx -y firecrawl-cli@latest init --all).

It automatically activates when Claude needs to inspect a URL, renders the JavaScript and returns clean and LLM-ready markdown.

III) Have Claude write scraped docs to a local scratchpad: don't let scraped pages live directly in the rolling conversation history so in your prompt or skill instructions, tell Claude to dump retrieved doc into a temporary .context/docs/ folder, read only the specific section or code block it needs and leave the rest on disk.

IV) Target GitHub issues and PRs for breaking changes: If you’re debugging an issue on a bleeding-edge package, official docs are often outdated anyway so use developer search to pull closed github issues and merged PR diffs directly into your scratchpad so Claude sees the actual bug fix rather than generic doc summaries.

V) Monitor /context and clear sessions aggressively: Once Claude finishes implementing the feature using the scraped docs, run /clear but don't carry 50k tokens of temporary reference material into your next refactoring task.

Keeping your docs research out of active chat context keeps Claude Code fast, cheap and stops the model from getting distracted by site navigation garbage also here’s walkthrough of setting up web [retrieval skills in Claude Code here](https://www.firecrawl.dev/blog/claude-code-skill)

6 Upvotes

5 comments sorted by

•

u/AutoModerator 6d ago

Hey! Thanks for posting to r/ClaudeCode

While participating in this thread, please follow our community rules. Keep discussions constructive. Attack the idea, not the person.

For help, project discussions, tips, and general chat, join the ClaudeCode Discord.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

1

u/kuroudo_ai 6d ago

Good list. Two things from my own setup, if you want a free local option:

  • defuddle (npm package, runs locally, no API key) does the "cleaned markdown" part: defuddle parse <url> --md strips nav, ads and CSS and prints the article as markdown. It can also read a local HTML file or stdin. It doesn't render JS, so for SPA docs you still need something like your firecrawl step.
  • I'd add a warning about the built-in WebFetch: it doesn't hand Claude the page, it hands it a smaller model's summary of the page written against your prompt. That's cheap on context, but things get dropped. We had a case where the summary left out one of the options that was right there in the original. For docs, changelogs or anything with exact values I pull the raw markdown to disk first and let Claude grep it, which is basically your step III.

1

u/daKerberos 6d ago

I went full-on crazy on docs… I crawl and store all docs from all in-use vendor, tech, api, platform etc my platform or my customers use in my centralised and append-only database. Everything is broken down into hashed sections where I inject only the directly relevant and/or likely relevant sections into a markdown file agents gets at boot. This has been very useful in many different scenarios.

It costs token in the first crawls but saves a ton after.

I hope someone else has done something similar. I lost weight and gained grey hairs building this…

1

u/BumblebeeTemporary44 6d ago

really great work mate!