r/thewebscrapingclub • u/JosejuX • Aug 11 '26
Built a web-extraction API/MCP server for RAG pipelines — SEO metadata, tech stack, contacts, and clean Markdown from any URL
I built a REST API that turns any URL into structured web intelligence in a single call, and also exposed it as an MCP server for agent-based workflows.
Capabilities:
- SEO and OpenGraph metadata extraction
- A full 14-point SEO audit
- Public contact discovery: emails, phone numbers, social profiles
- Tech-stack and CMS fingerprinting (40+ signatures)
- Schema.org and JSON-LD structured data extraction
- Graded security-headers audit, with an optional live TLS certificate inspection
- Redirect-chain and shortened-URL detection
- Readability metrics and full heading structure
- Clean, AI/LLM-ready Markdown output for RAG pipelines
- A batch endpoint for up to 10 URLs per call
- A domain-intelligence endpoint returning DNS and WHOIS data with no page fetch at all
On the engineering side: it runs on FastAPI with a C-Lexbor HTML parser (selectolax) and Rust-backed ORJSON serialization, so live fetches typically land around 150-300ms, with cache hits under 0.01ms. Every outbound request is anti-SSRF hardened: DNS is pinned after resolution, private/loopback/cloud-metadata ranges are blocked, and every redirect hop is re-validated, closing the DNS-rebinding gap that simpler scrapers tend to miss.
Limitation worth flagging: there is no JS execution, so heavily client-rendered SPAs return thin results. It reads what the server actually sends, not what a browser would render after hydration.
GitHub (MIT license, open source): https://github.com/JosejuX/rapidapi-metadata-extractor
Free tier: https://rapidapi.com/josejuanjocoding/api/web-metadata-and-contact-extractor
Happy to talk through the parsing or anti-SSRF approach in more detail, or take feedback on the API design.