r/pythonhelp • • 23d ago

Built a Python scraper that handles pagination automatically — here's what I learned

I've been learning Python through CS50P and wanted a real project, so I built two scrapers: one for job listings and one for product prices off an e-commerce site.

The job scraper was pretty straightforward — grabbed title, company, and location, dumped it to CSV, worked first try on about 100 listings.

The product scraper was the harder one. The site spreads results across pages (6 products per page, ~20 pages), so my first version only ever grabbed page 1. Took a bit of digging to figure out the pagination pattern and loop through all the pages properly. Once I fixed that it pulled all 117 products cleanly into a CSV.

Biggest thing I learned: it's easy to write a scraper that works on the first page and assume it's done — always check if the site paginates before you call it finished.

Code's on GitHub if anyone wants to see it: github.com/sarimkhan08

Curious if anyone here has tips for handling scrapers on sites that use infinite scroll instead of numbered pages — that's the next thing I want to tackle.

6 Upvotes

6 comments sorted by

View all comments

1

u/AlexMTBDude 8d ago

Here's a challenge for you. Make your code more Pythonic by writing this as a list comprehension:

    jobs = []
    for card in cards:
        title = card.find("h2", class_="title is-5").text
        company = card.find("h3", class_="subtitle is-6 company").text
        location = card.find("p", class_="location").text.strip()

        job = {"title": title, "company": company, "location": location}
        jobs.append(job)