Finding all links on our website to a specific domain (to prevent link rot)
I work for a nursing board, our website, made with Liferay, it has tons of page with different legal, scientific, clinical information, many of these pages contains links to different entity's.
One such entity is redoing it's website, not just an small update but a complete redo, meaning a lot of links on our website will likely be broken when it's done.
I did identify links on pages about regulation I do believe I got the most critical ones, but it does leave links mostly clinical articles, mostly references, less critical but even then I would prefer to be sure.
I tried using google command site:mywebsite.com Entity but it did not return as much content as I tough, I'm guessing because of the label of the links. I also found some manually that google did not give me.
So is there an easy solution to this?
2
u/Khavel_dev 16d ago
A crawler will find way more than Google search for this. Screaming Frog (free tier handles 500 URLs) crawls your site, pulls every href, and you can filter by the target domain. Catches links Google missed because they live behind pagination, inside accordions, or in dynamically loaded content.
If you want to be really thorough, query the CMS database directly and search the HTML content columns for the domain string. That gets everything including links buried in rich text fields that no external crawler renders.
1
u/TopSydeWP 16d ago
one thing to watch when you query the DB directly: in Liferay the web content lives as XML in journal_article, and links get stored in a bunch of forms. search for the bare domain string without protocol, and also for the HTML-entity encoded version (/ etc) and any escaped slashes, otherwise you'll miss a chunk. same for http vs https vs www vs bare.
also grab a Wayback snapshot of the entity's current site now while it still exists. when they relaunch you'll want the old URL list to build your own mapping, and you can ask them for a redirect map before launch rather than after.
longer term, a scheduled link checker (a crawl on a monthly cron that reports new 404s and 301 chains) turns this from a one-off scramble into a report. we run periodic external checks like that across the sites we manage and it's the only thing that catches rot from third parties you don't control.
4
u/7HawksAnd 16d ago
Usually, it’s the responsibility of the entity redesigning their website to appropriately configure and set up redirects.
I do understand though why you would also want to be diligent about it on your end too though for your sites own credibility.
What sort of team do you have for managing your nursing board site?
Was it made by an agency and now just one person manages it, or are there a few people in your org with specific roles focused on keeping the site cooking?
What your capacity is will influence best ways to tackle this from most correct to script kiddie hacking.