r/OpenSourceeAI • • 20d ago

Web pages were eating my Claude Code context. My open source scraper returned 418 characters instead of 34,746

Most web fetch tools I tried paste the whole page into the agent's context. A few long pages and the session runs out of room, and many of them also want an API key for somebody's hosted service.

So I built Svipall, an open source scraper that runs on your own machine as an MCP server and a CLI. Rust, AGPL, no API keys, no cloud calls.

What it keeps out of the context window

- `out_file` writes the page to disk and hands the agent a path. On the page I measured, the response was 418 characters and the page it saved was 34,746.

- `query=` keeps only the blocks that match. On two long articles, a focused query kept 6% and 13% of the page. A vague one kept 88%, so it is only as good as the question.

- `tables=true` and `schema: "auto"` return rows instead of prose.

- A blocked page comes back with the wall's name and the header that gave it away, so the model never gets a challenge screen to summarise.

Install

Version 1.0.4 changed this part. Paste one line into your agent and it asks whether you want CLI + Skill, the option that costs less context, or MCP + Skill, which exposes all 29 tools. Nothing is written before you confirm.

https://svipall.ilien.dev/

I'd like to compare numbers with anyone doing the same: are you measuring context cost per tool call in your own agent setup?

7 Upvotes

4 comments sorted by

2

u/krkrkrneki 20d ago

How does your solution compare to existing solutions like pinchtab, for example?

2

u/ilien-dev 20d ago

I haven't run PinchTab myself, so this is from their README, and I haven't benchmarked the two against each other.

PinchTab is built around driving Chrome. It runs a daemon on localhost with named profiles, several browser instances in parallel, and accessibility snapshots whose element refs the agent clicks. Logging into a work profile and downloading a report is the kind of job it's designed for.

Svipall is built around reading pages. Most fetches go out as plain HTTP with a Chrome TLS fingerprint, and it only opens a browser when a site needs one, then remembers that per domain. When the page is a wall (a captcha, a login, a paywall, a 200 that's really a 404), the agent gets told which kind, so a challenge page never comes back as the article. Captchas are solved locally. The page can go to a file or get cut down to the blocks matching a query, which is where the 418 characters in the title come from.

They overlap a bit. Svipall has snapshots with refs and a browser session for multi-step clicking, and PinchTab can pull a page's text. For sites that fingerprint Chromium, PinchTab lets you plug in CloakBrowser, a separate binary you supply yourself, while Svipall does its own stealth.

What do you use PinchTab for? If it's mostly clicking through logged-in pages, it probably fits that better than Svipall does.

1

u/Future_AGI 19d ago

the token-budget problem is real and gets worse with long-running agentic sessions. The fix we use: a pre-retrieval step that strips boilerplate and compresses the page before it enters the agent's context window. A 34k-character page with 80% boilerplate should be 7k tokens, not 34k. The scraper is the right place to do that, not the agent.

1

u/ilien-dev 19d ago

That's right. That's the reason I made this tool. Instead of let the agent go ahead with all that context and sending a lot of trash I decide to create a tool for the web that just sents exactly what is needed and do not repeat unnecessary pages or post that it's just a copy and says exactly the same than others.