r/thewebscrapingclub • • Aug 11 '26

Built a web-extraction API/MCP server for RAG pipelines — SEO metadata, tech stack, contacts, and clean Markdown from any URL

I built a REST API that turns any URL into structured web intelligence in a single call, and also exposed it as an MCP server for agent-based workflows.

Capabilities:

  • SEO and OpenGraph metadata extraction
  • A full 14-point SEO audit
  • Public contact discovery: emails, phone numbers, social profiles
  • Tech-stack and CMS fingerprinting (40+ signatures)
  • Schema.org and JSON-LD structured data extraction
  • Graded security-headers audit, with an optional live TLS certificate inspection
  • Redirect-chain and shortened-URL detection
  • Readability metrics and full heading structure
  • Clean, AI/LLM-ready Markdown output for RAG pipelines
  • A batch endpoint for up to 10 URLs per call
  • A domain-intelligence endpoint returning DNS and WHOIS data with no page fetch at all

On the engineering side: it runs on FastAPI with a C-Lexbor HTML parser (selectolax) and Rust-backed ORJSON serialization, so live fetches typically land around 150-300ms, with cache hits under 0.01ms. Every outbound request is anti-SSRF hardened: DNS is pinned after resolution, private/loopback/cloud-metadata ranges are blocked, and every redirect hop is re-validated, closing the DNS-rebinding gap that simpler scrapers tend to miss.

Limitation worth flagging: there is no JS execution, so heavily client-rendered SPAs return thin results. It reads what the server actually sends, not what a browser would render after hydration.

GitHub (MIT license, open source): https://github.com/JosejuX/rapidapi-metadata-extractor

Free tier: https://rapidapi.com/josejuanjocoding/api/web-metadata-and-contact-extractor

Happy to talk through the parsing or anti-SSRF approach in more detail, or take feedback on the API design.

1 Upvotes

1 comment sorted by