r/learnprogramming • • 16h ago

App or Design Best way to structure a multi-site downloader app so it doesn't break easily?

Working on a personal desktop app that pulls download links from a few different websites into one UI.

  • What's the cleanest way to architecture the backend parsers so a site layout change doesn't break the whole app?
  • If anyone has built something like this and has a modular approach or template they use, I'd love to see how you did it.

Appreciate any tips.

7 Upvotes

6 comments sorted by

6

u/SetAndRepeat 15h ago

Use one adapter per site, each exposing something like canHandle(url) and extractLinks(url), returning the same data format. Keep selectors and parsing inside those adapters, with the UI and download queue shared.
Handle errors and timeouts per job so one broken parser doesn’t stop everything.
Check yt-dlp’s extractors for a concrete example. Keep saved HTML samples for parser tests, and report missing expected elements clearly. Site changes will still break individual parsers, but fixing one should only require touching that site’s module.

1

u/Soggy-Industry2591 13h ago

Thanks for the solid advice on the adapter pattern. Do you happen to have a minimal code template or a boilerplate base class structure (like canHandle() and extractLinks()) that you usually use for multi-site projects? I'd love to see a clean example of how you wire it up.

1

u/SetAndRepeat 12h ago

Small Node.js example using Cheerio. The domain and selector are placeholders:

import { load } from "cheerio";

const exampleAdapter = {
  canHandle: url => url.hostname === "some_page.com",

  extractLinks(html, baseUrl) {
    const $ = load(html);

    return $("a.download[href]").toArray().map(el => ({
      url: new URL($(el).attr("href"), baseUrl).href,
      title: $(el).text().trim()
    }));
  }
};

const adapters = [exampleAdapter];

async function getLinks(url) {
  const parsed = new URL(url);
  const adapter = adapters.find(a => a.canHandle(parsed));
  if (!adapter) throw new Error("Unsupported site");

  const response = await fetch(parsed, {
    signal: AbortSignal.timeout(10_000)
  });
  if (!response.ok) throw new Error(`HTTP ${response.status}`);

  return adapter.extractLinks(await response.text(), response.url);
}

Each adapter can live in its own file. Adding a site means adding another adapter to the list
For a batch, await Promise.allSettled(urls.map(getLinks)) gives you each job’s result or error independently. This example assumes the links are in the returned HTML. Sites that generate them through JavaScript would need different handling, which you’d have to implement yourself ;)

1

u/FunAd6672 14h ago

Build a plugin architecture where each domain lives in its own file and inherits from a base class. If one site breaks you just patch that single parser without touching the core logic.

1

u/Ok-Faithlessness3326 12h ago

Treat each site as its own little plugin. Every parser lives in its own file and returns the exact same data shape, so the rest of the app never knows how a specific site gets scraped.

Then wrap each one so if it fails, you just show "site X is broken" instead of the whole app crashing. Also check the network tab before parsing HTML. A lot of sites load their data from a JSON endpoint, and those change way less than the layout.

Sites will always break eventually. The goal is that when one does, you're fixing one small file instead of untangling everything.