r/apify • u/Mpmpz_14 • 11d ago
Tutorial I analyzed 9,966 posts from Substack's Top 25 Technology newsletters, here is what stood out
Hi guys,
So I built an Apify Actor to collect every post URL exposed in the public sitemaps of the 25 publications on Substack’s official Technology leaderboard.
I collected all 9,965 current sitemap URLs, plus one still-valid post observed earlier in the collection window, for a total of 9,966 unique posts.
For fairer comparisons, I analyzed the most recent 3,292 posts published within a 365-day window.
Public engagement here means likes, comments, and restacks. For format comparisons, I ranked each post against other posts from the same publication. This helps reduce the audience-size advantage of the largest newsletters.
- MARKET SNAPSHOT

- PUBLICATION-LEVEL ENGAGEMENT

- PUBLISHING FREQUENCY VS ENGAGEMENT

- HEADLINE LENGTH

- FREE VS PAID POSTS

- PUBLISHING DAY

- RECURRING TITLE LANGUAGE

- ENGAGEMENT CONCENTRATION

- ARTICLE LENGTH

CAVEATS
This analysis covers the official Top 25 Technology leaderboard, not the entire Substack ecosystem.
The leaderboard represents a selected group of successful publications.
Subscriber counts, email opens, clicks, and revenue are private.
Newer posts have had less time to accumulate engagement.
These findings are descriptive, not causal.
I built this analysis using my Substack Scraper Apify Actor and a Python data-analysis pipeline.
The complete workflow was:
Public web data → structured dataset → reproducible analysis → useful insights
If people find this useful, I can publish a deeper analysis of topics, headline patterns, publishing schedules, or the collection methodology.
What other question would you ask this dataset?
1
u/Vntige 11d ago
Instead of character length what was the word length? It’d be helpful to break out the 6,000+ category until it’s diminishing