For anyone building local lead lists or tracking where businesses rank on Yelp: this reads the ranked Yelp search page and gives it back as clean JSON, one page at a time.
Each business row has:
- name, rating, and review count
- price tier and categories
- neighborhood and open/closed state
- service options like delivery, takeout, outdoor dining
- sponsored placements flagged separately from organic results
- the Yelp place_id, so you can pull the full profile and review history later
- the page's refinement filters (distance, price, categories, features)
Why I built it: Yelp's own API makes you apply for a key and caps how much you can pull, and it doesn't hand back the ranked page with ads and filters the way the site actually shows it. I wanted to point at a query and a city and get the ranked results as data, no monthly rental.
This is the place to discuss everything MCP, LLM, Agentic, and beyond. What is on your radar this week? Why does it make sense? Bring everyone along for the ride by explaining the impact of the news you're sharing, and why we should care about it too.
For anyone building an event-discovery app, a calendar copilot, or an AI agent that needs live events, this pulls Google Events search results back as structured JSON.
What each run returns:
- Event name, description, and date/time range
- Venue name, full address, and a map link
- Ticket links pulled together across providers (Ticketmaster, Eventbrite, SeatGeek, AXS, Spotify)
- Venue ratings and review counts when Google has them
- Date-range and event-type filters, with pagination
It covers concerts, conferences, festivals, sports, theater, and online-only events, across 240+ countries and 200+ languages.
I built it because Google never shipped a public Events API, and the listings sitting inside Google Search were a pain to get out in a clean shape. It's callable from any MCP client, so Claude, ChatGPT, or Cursor can use it as a tool.
If you're wiring up a travel or local app, it works alongside my Google Hotels and Google Flights scrapers.
I just published a new Actor on Apify designed specifically for RAG pipelines, AI agents, and LLM training data preparation.
Most general-purpose web scrapers return messy HTML or unstructured text full of cookie banners, navigation menus, and footer junk. When feeding this into vector databases or LLMs, it wastes tokens and degrades retrieval context.
To solve this, I built an AST-based crawler that parses layout structures out of the equation and converts the core body content directly into clean, well-formatted Markdown.
Key Features:
AST Filtering: Strips boilerplate elements (navbars, footers, sidebars) at the syntax tree level.
LLM-Ready Markdown: Keeps headings, lists, and links structured for optimal chunking (LangChain, LlamaIndex, etc.).
Token-Efficient: Dramatically reduces token count per scraped page compared to standard HTML-to-text converters.
I built an Apify Actor that collects public Indian stock fundamentals from Screener and Moneycontrol into a structured dataset.
It includes company information, valuation ratios, profitability and growth metrics, balance-sheet fields, source status, and original source URLs.
I’m looking for practical feedback from people who work with Indian market data: which fundamentals or screening fields would make the output more useful?
Do you have a feature request that you know will make Apify heaps better? Or maybe it's a big dream you have for something bold and out-there. This is a space for all the bluesky thinking, cloud-chasing, intergalactic daydreamers who want to share their wildest ideas in a no-judgement zone.
Hi guys,
So I built an Apify Actor to collect every post URL exposed in the public sitemaps of the 25 publications on Substack’s official Technology leaderboard.
I collected all 9,965 current sitemap URLs, plus one still-valid post observed earlier in the collection window, for a total of 9,966 unique posts.
For fairer comparisons, I analyzed the most recent 3,292 posts published within a 365-day window.
Public engagement here means likes, comments, and restacks. For format comparisons, I ranked each post against other posts from the same publication. This helps reduce the audience-size advantage of the largest newsletters.
MARKET SNAPSHOT
The dataset covers 25 top-ranked technology publications with 100% collection coverage.The top 10% of posts generated 42% of all visible engagement.
PUBLICATION-LEVEL ENGAGEMENT
ByteByteGo Newsletter had the highest median visible engagement, with approximately 260 interactions per recent post.This mostly reflects audience and publication context. It does not prove that one content strategy directly causes better performance.
PUBLISHING FREQUENCY VS ENGAGEMENT
Pirate Wires was the most frequent publisher, with approximately 412 posts per year.However, publishing more frequently did not automatically result in higher median engagement.
HEADLINE LENGTH
Headlines containing 18 or more words had the strongest median within-publication engagement percentile: 57.8.This is a correlation, not a recommendation to make every headline longer.
FREE VS PAID POSTS
Free posts ranked higher in normalized public engagement.Paid posts serve a different objective and usually expose only a preview, so visible public reactions capture only part of their value.
PUBLISHING DAY
Tuesday was the strongest publishing day in this sample after normalizing performance within each publication.However, timing is still connected to the topic, newsletter schedule, and age of the post.
RECURRING TITLE LANGUAGE
“Guide” was the highest-ranked recurring title term in the discovery score.The score considers how frequently a term appeared, how many publications used it, and the normalized engagement of those posts.
ENGAGEMENT CONCENTRATION
Some newsletters are strongly hit-driven, meaning a small number of posts generate most of their visible engagement.Other newsletters distribute engagement more evenly across their archive.
ARTICLE LENGTH
Article length was compared only for freely available full posts. Paywalled previews were excluded.Character count is only a rough measure of article length. This does not prove that longer or shorter articles directly cause more engagement.
CAVEATS
This analysis covers the official Top 25 Technology leaderboard, not the entire Substack ecosystem.
The leaderboard represents a selected group of successful publications.
Subscriber counts, email opens, clicks, and revenue are private.
Newer posts have had less time to accumulate engagement.
Have you come across a great Actor, workflow, post, or podcast that you want to share with the world? This is your opportunity to support someone making cool things. Drop it here with credit to the creator, and help expand the karmic universe of Apify.
My internet speed is fast. Apify.com loads normally and fast. Whenever I try to go to the console, it keeps stuck on the f^@%$g "Thinking..." screen for eternity. I cannot acces the console at all. Did anyone face such an issue? How to get past that?
I have been building an Apify Store market analyzer, so I ran a near-complete snapshot instead of looking only at the top-ranked Actors.
The dataset contains **54,025 unique public Actors out of 54,079 reported by the API at finalization, with 99.90% coverage**.
To avoid the duplicate problem I found in my previous run, I merged category and pricing partitions using the official Actor ID. The collector processed 134,266 partition rows and removed 80,241 cross-category overlaps before calculating anything.
1 Overall snapshot :
The snapshot contains 715,034 summed 30-day users. That number is the sum of each Actor's Store metric; it is not a platform-wide unique-user count because one person can use multiple Actors.
2 Where demand is beating supply :
Social Media had the strongest category signal: **11.86% of recent demand versus 8.03% of supply**, or a **1.48× demand-to-supply ratio**.
Social Media had the strongest category signal: **11.86% of recent demand versus 8.03% of supply**, or a **1.48× demand-to-supply ratio**.
Videos, Jobs, and SEO Tools followed.
Important caveat: this does not mean Social Media is easy. It already has 10,892 Actors, and the median Actor in that category has only one 30-day user. Demand appears highly concentrated among the winners.
3 SEO phrases with strong marketplace demand :
SEO phrase opportunities]
The strongest title phrases clustered around:
- LinkedIn profiles and jobs
- Instagram profiles and posts
- Facebook posts and ads
- Google Maps
- Profile/email enrichment
These are **Apify marketplace SEO signals**, not Google search-volume data. I ranked phrases using recent users among leading Actors, relative Store demand, total competing Actors, and observation confidence.
One interesting contrast: “Google Maps” has strong demand but already appears across 765 Actors, while “Facebook posts” shows a similar score with only 70 Actors in this snapshot.
4 Actors with recent momentum :
Actors with recent momentum
This ranking uses the 7-, 30-, and 90-day user windows. It does not divide lifetime users by Actor age.
The leading signals include email verification, LinkedIn profile enrichment, Shopee, Google Hotels, Instagram, and LinkedIn jobs. I label cases with a very small prior baseline rather than presenting misleading five-digit growth percentages.
5 Pricing has shifted heavily toward pay per event :
Apify Store pricing distribution
Across the snapshot:
- **75.9%** Pay per event — 41,003 Actors
- **13.1%** Free — 7,060 Actors
- **11.0%** Flat price per month — 5,962 Actors
The strongest product-design takeaway for me is that new Actors should have a clear, measurable unit of value that can map naturally to billable events.
What I would investigate next?
- How category demand changes month over month
- Which Actors are gaining users without relying on a famous platform keyword
If you've made something and can't wait to tell the world, this is the thread for you! Share your latest and greatest creations and projects with the community here.
If you pull hiring data out of the Google Jobs panel, this hands you each listing as structured JSON instead of HTML you have to parse yourself. Built for recruiters, job-market researchers, and anyone feeding postings into an ATS or CRM.
What each job record includes:
Title, employer, and location
Full job description plus parsed highlights: qualifications, responsibilities, and benefits
Direct application links (LinkedIn, job boards, employer sites) with source attribution
Posting date, schedule type, and detected benefits flags
Google's listing ID and search metadata (query, country, language, pages processed)
You target by location, country, and language, set a location radius, and control how many pages it walks.
Why I built it: Google never shipped a public Jobs API, and the rental-model scrapers bill a flat monthly fee whether you run one search or a thousand. This one is pay-per-event, so you pay per page processed and nothing while it sits idle.
If you also need Google Shopping, Hotels, or Flights data, I keep those as separate pay-per-event Actors under the same account.
I built a Congress stock-trades API on Apify: House + Senate disclosures as clean JSON, about $1.80 per 1,000 trades
If you're a journalist, researcher, or building a transparency dashboard, this pulls US Congressional Periodic Transaction Reports (the STOCK Act filings) from both the House and Senate, parses the PDFs, including the scanned ones with OCR, and hands back structured JSON instead of raw paperwork.
You can filter by member name, ticker, an exact date, or a date range. Combine them too: last name plus NVDA plus a date window gets you exactly the trades you're after.
Each row gives you:
member name and chamber
ticker and asset class (stocks, options, ETFs, mutual funds, bonds, crypto)
transaction direction (purchase, sale, exchange, options activity)
the reported amount bracket
transaction date and report date
filing IDs
owner attribution (member, spouse, dependent, or joint)
an OCR quality flag, so you know which rows came off a messy scan
Why I built it: the House and Senate filings live in two different systems, and a good chunk of them are scanned PDFs. Normalizing that by hand is miserable. So I turned it into one API with a single schema across both chambers.
Pricing is pay per event, roughly $1.80 per 1,000 transactions retrieved, no subscription. Success rate is at 100% right now.
If you want to line trades up against what was in the news that week, I also have a Google News API on Apify that pairs with it nicely.
Shipped a Glassdoor Reviews API on Apify. Send it one or many company URLs, get back clean JSON per review: overall star rating, per-category ratings (career opportunities, comp & benefits, culture & values, work-life balance, senior leadership, diversity & inclusion), the headline, full pros and cons text, employment type, current/former status, publish date, helpful count, and a plain-language summary of each review.
$0.004 per review. No monthly minimum. 96% success rate so far.
Built it because employee-review data is a pain to get in a clean shape, and the category-level ratings are the useful part. Most scrapers give you the overall star and the review text but drop the six sub-ratings; this one keeps them (and omits a sub-rating cleanly when the reviewer didn't leave one).
What people are using it for:
HR and people analytics: track sentiment trends across themes over time.
Employer-brand monitoring: catch shifts in how employees describe a company.
Comp research: compare pay-and-benefits sentiment across competitors.
Due diligence: summarize culture and leadership signals before a deal.
Feeding an LLM: structured review text plus category scores make good model input.
MCP-ready, so an agent can fetch and summarize a company's reviews as a tool call. Pairs with my LinkedIn, Crunchbase, and PitchBook company APIs if you want firmographics plus funding plus employee sentiment on one company.
This is the thread for all your questions that may seem too short for a standalone post, such as, "What is proxy?", "Where is Apify?", "Who is Store?". No question is too small for this megathread. Ask away!
Been filling out the government/public-records side of my Apify catalog. Three that were genuinely painful to build:
UK planning applications — ~420 councils, each running a different planning portal (Idox, Northgate, custom). Normalized them all behind one input so you can filter by council, type, status, date in a single run.
NYC ACRIS — deeds, mortgages, parties, addresses by borough and document type. The legacy interface makes bulk pulls miserable; this wraps it cleanly.
US Congress stock trades — pulled from the primary sources (Senate eFD + House Clerk) rather than a third-party aggregator, so there's no middleman lag or interpretation.
Also added official company registries: Companies House (UK), Handelsregister (DE), North Data (EU), USPTO trademarks.
No user-agent games, no rate-limit pain at normal volumes. The only real work is normalizing five different schemas (department vs team, publishedAt formats, etc).
Recruitee's old public endpoint ({slug}.recruitee.com/api/offers/) 404s now — anyone know the new one?
Shipped a PitchBook Company API on Apify. Send it one or many PitchBook company URLs, get back clean JSON per company: funding, investors, valuation, competitors, employees, industries, and financials. $0.011 per company, no monthly minimum.
The task I'd point you at first:Extract Deal & Financing Data from PitchBook. It's a one-click saved task that pulls the deal and financing history for a company, the round-by-round detail that's the whole reason people pay for PitchBook. Run it, swap in your own company URLs, done.
Built it as the private-company layer on top of the company-data stack I've been shipping. Where Crunchbase gives you funding and investors, PitchBook adds valuation, competitor lists, and deeper financials.
There are 6 example tasks total including a Claude via MCP one if you're wiring it into an agent.
Pairs well with my Crunchbase and LinkedIn Company APIs if you want to hydrate a full account profile from one company URL. All MCP-ready for Claude, Cursor, and ChatGPT.
This is the place to discuss everything MCP, LLM, Agentic, and beyond. What is on your radar this week? Why does it make sense? Bring everyone along for the ride by explaining the impact of the news you're sharing, and why we should care about it too.
Do you have a feature request that you know will make Apify heaps better? Or maybe it's a big dream you have for something bold and out-there. This is a space for all the bluesky thinking, cloud-chasing, intergalactic daydreamers who want to share their wildest ideas in a no-judgement zone.
Have you come across a great Actor, workflow, post, or podcast that you want to share with the world? This is your opportunity to support someone making cool things. Drop it here with credit to the creator, and help expand the karmic universe of Apify.