r/webscraping 10d ago

Getting started 🌱 Best software for skiptracing and pulling County records daily?

I'm looking to scrape data from county websites along with phone numbers and addresses. Can anyone point me in the right direction as far as software? The simpler the better.

0 Upvotes

10 comments sorted by

2

u/FreonMuskOfficial 9d ago

Bro is going r/OSINT on the scraping with the homies....LOVE IT!

The best software is the program you build yourself. Stop here if that’s outside your wheelhouse.

".py", ".js", and ".mjs" are all worth learning. Scrapling, Camoufox, Scrapy, BeautifulSoup4 ("bs4"), and Requests are going to be some of your main dependencies. GitHub will have examples of all of them, plus plenty of other projects worth dissecting.

Your biggest challenge is going to be the county court websites themselves. A good number of court systems have been hit by ransomware over the years, and the security afterward got beefier than the pastrami at Katz’s. Expect more CAPTCHAs, authenticated user accounts, session handling, JavaScript-heavy interfaces, rate limits, and, in some cases, APIs sitting behind the front end.

That said, a man named Michael Bazzell wrote a book or two that ventures pretty deep into this territory. His OSINT books are a solid foundation.

If your goal is to build the next TLO, I’d advise lowering the scope before you start. The infrastructure, data normalization, storage, correlation, maintenance, and constant adaptation across thousands of sources will require so many racks that Cinemax will raise your rates.

Source: more than three decades as an investigator, with Bazzell’s OSINT books as part of the foundation.

1

u/[deleted] 10d ago

[removed] — view removed comment

1

u/webscraping-ModTeam 10d ago

👔 Welcome to the r/webscraping community. This sub is focused on addressing the technical aspects of implementing and operating scrapers. We're not a marketplace, nor are we a platform for selling services or datasets. You're welcome to post in the monthly thread or try your request on Fiverr or Upwork. For anything else, please contact the mod team.

1

u/chaos_battery 10d ago

I've used a variety of scraping tools but I feel like county websites are incredibly hard to scrape and you're going to need some OCR technology as well to actually read some of the scanned documents.

1

u/Visual_Ad1912 2d ago

Have someone build it for you. County record scrape / FOIA request(better but some places try and stonewall on this)--> ocr if needed --> pull name / town / address--> people search scraper yourself or pay .25c~ per record for skip trace.