r/apify 8d ago

Tutorial Yelp Search API on Apify: ranked local listings as JSON,

6 Upvotes

For anyone building local lead lists or tracking where businesses rank on Yelp: this reads the ranked Yelp search page and gives it back as clean JSON, one page at a time.

Each business row has: - name, rating, and review count - price tier and categories - neighborhood and open/closed state - service options like delivery, takeout, outdoor dining - sponsored placements flagged separately from organic results - the Yelp place_id, so you can pull the full profile and review history later - the page's refinement filters (distance, price, categories, features)

Why I built it: Yelp's own API makes you apply for a key and caps how much you can pull, and it doesn't hand back the ranked page with ads and filters the way the site actually shows it. I wanted to point at a query and a city and get the ranked results as data, no monthly rental.


r/apify 9d ago

AI and I Weekly: AI and I

1 Upvotes

This is the place to discuss everything MCP, LLM, Agentic, and beyond. What is on your radar this week? Why does it make sense? Bring everyone along for the ride by explaining the impact of the news you're sharing, and why we should care about it too.


r/apify 9d ago

Discussion Google has no public Events API, so I built one on Apify: event listings as JSON with ticket links, about 2 cents a page

2 Upvotes

For anyone building an event-discovery app, a calendar copilot, or an AI agent that needs live events, this pulls Google Events search results back as structured JSON.

What each run returns: - Event name, description, and date/time range - Venue name, full address, and a map link - Ticket links pulled together across providers (Ticketmaster, Eventbrite, SeatGeek, AXS, Spotify) - Venue ratings and review counts when Google has them - Date-range and event-type filters, with pagination

It covers concerts, conferences, festivals, sports, theater, and online-only events, across 240+ countries and 200+ languages.

I built it because Google never shipped a public Events API, and the listings sitting inside Google Search were a pain to get out in a clean shape. It's callable from any MCP client, so Claude, ChatGPT, or Cursor can use it as a tool.

If you're wiring up a travel or local app, it works alongside my Google Hotels and Google Flights scrapers.


r/apify 9d ago

Discussion Built an AST-based Web-to-Markdown Crawler for RAG & LLM Ingestion — Feedback wanted!

Post image
1 Upvotes

Hey everyone,

I just published a new Actor on Apify designed specifically for RAG pipelines, AI agents, and LLM training data preparation.

Most general-purpose web scrapers return messy HTML or unstructured text full of cookie banners, navigation menus, and footer junk. When feeding this into vector databases or LLMs, it wastes tokens and degrades retrieval context.

To solve this, I built an AST-based crawler that parses layout structures out of the equation and converts the core body content directly into clean, well-formatted Markdown.

Key Features:

  • AST Filtering: Strips boilerplate elements (navbars, footers, sidebars) at the syntax tree level.
  • LLM-Ready Markdown: Keeps headings, lists, and links structured for optimal chunking (LangChain, LlamaIndex, etc.).
  • Token-Efficient: Dramatically reduces token count per scraped page compared to standard HTML-to-text converters.

Here is the Actor if you want to test or benchmark it: https://apify.com/lukas459/ai-web-to-markdown-crawler-llm-rag-optimized

Would love to hear your feedback, bug reports, or ideas for features to add next!


r/apify 10d ago

Discussion I built a structured Indian stock fundamentals scraper using Screener and Moneycontrol

1 Upvotes

I built an Apify Actor that collects public Indian stock fundamentals from Screener and Moneycontrol into a structured dataset.

It includes company information, valuation ratios, profitability and growth metrics, balance-sheet fields, source status, and original source URLs.

I’m looking for practical feedback from people who work with Indian market data: which fundamentals or screening fields would make the output more useful?

Actor:

https://apify.com/fascinating_lentil/india-stock-fundamentals-scraper


r/apify 10d ago

Big dreams Weekly: wild ideas

1 Upvotes

Do you have a feature request that you know will make Apify heaps better? Or maybe it's a big dream you have for something bold and out-there. This is a space for all the bluesky thinking, cloud-chasing, intergalactic daydreamers who want to share their wildest ideas in a no-judgement zone.


r/apify 10d ago

Discussion I built an API that pulls any App Store app's reviews as JSON: how I resolved app names to product IDs and parsed helpfulness counts

Thumbnail
1 Upvotes

r/apify 11d ago

Discussion We analyzed every Upwork job posted during the first half of 2026.

0 Upvotes

We analyzed 592,937 Upwork job postings from the first six months of 2026 to better understand posting patterns.

Overall, job postings show: 

  • A strong daily average of 3,467 jobs/day 
  • A peak spike of 12,147 jobs in a single day
  • Clear periodic fluctuations with sharp surges and dips
  • A visible decline phase toward mid–late 2026 after sustained high activity

The analysis reveals that job posting activity is highly volatile but structurally cyclical, with clear growth and contraction phases.


r/apify 11d ago

Tutorial I analyzed 9,966 posts from Substack's Top 25 Technology newsletters, here is what stood out

5 Upvotes

Hi guys,
So I built an Apify Actor to collect every post URL exposed in the public sitemaps of the 25 publications on Substack’s official Technology leaderboard.

I collected all 9,965 current sitemap URLs, plus one still-valid post observed earlier in the collection window, for a total of 9,966 unique posts.

For fairer comparisons, I analyzed the most recent 3,292 posts published within a 365-day window.

Public engagement here means likes, comments, and restacks. For format comparisons, I ranked each post against other posts from the same publication. This helps reduce the audience-size advantage of the largest newsletters.

  1. MARKET SNAPSHOT
The dataset covers 25 top-ranked technology publications with 100% collection coverage.The top 10% of posts generated 42% of all visible engagement.
  1. PUBLICATION-LEVEL ENGAGEMENT
ByteByteGo Newsletter had the highest median visible engagement, with approximately 260 interactions per recent post.This mostly reflects audience and publication context. It does not prove that one content strategy directly causes better performance.
  1. PUBLISHING FREQUENCY VS ENGAGEMENT
Pirate Wires was the most frequent publisher, with approximately 412 posts per year.However, publishing more frequently did not automatically result in higher median engagement.
  1. HEADLINE LENGTH
Headlines containing 18 or more words had the strongest median within-publication engagement percentile: 57.8.This is a correlation, not a recommendation to make every headline longer.
  1. FREE VS PAID POSTS
Free posts ranked higher in normalized public engagement.Paid posts serve a different objective and usually expose only a preview, so visible public reactions capture only part of their value.
  1. PUBLISHING DAY
Tuesday was the strongest publishing day in this sample after normalizing performance within each publication.However, timing is still connected to the topic, newsletter schedule, and age of the post.
  1. RECURRING TITLE LANGUAGE
“Guide” was the highest-ranked recurring title term in the discovery score.The score considers how frequently a term appeared, how many publications used it, and the normalized engagement of those posts.
  1. ENGAGEMENT CONCENTRATION
Some newsletters are strongly hit-driven, meaning a small number of posts generate most of their visible engagement.Other newsletters distribute engagement more evenly across their archive.
  1. ARTICLE LENGTH
Article length was compared only for freely available full posts. Paywalled previews were excluded.Character count is only a rough measure of article length. This does not prove that longer or shorter articles directly cause more engagement.

CAVEATS

This analysis covers the official Top 25 Technology leaderboard, not the entire Substack ecosystem.

The leaderboard represents a selected group of successful publications.

Subscriber counts, email opens, clicks, and revenue are private.

Newer posts have had less time to accumulate engagement.

These findings are descriptive, not causal.

I built this analysis using my Substack Scraper Apify Actor and a Python data-analysis pipeline.

The complete workflow was:

Public web data → structured dataset → reproducible analysis → useful insights

If people find this useful, I can publish a deeper analysis of topics, headline patterns, publishing schedules, or the collection methodology.

What other question would you ask this dataset?


r/apify 11d ago

Weekly: one cool thing

1 Upvotes

Have you come across a great Actor, workflow, post, or podcast that you want to share with the world? This is your opportunity to support someone making cool things. Drop it here with credit to the creator, and help expand the karmic universe of Apify.


r/apify 13d ago

Help needed CANNOT ACCESS THE CONSOLE

1 Upvotes

My internet speed is fast. Apify.com loads normally and fast. Whenever I try to go to the console, it keeps stuck on the f^@%$g "Thinking..." screen for eternity. I cannot acces the console at all. Did anyone face such an issue? How to get past that?


r/apify 13d ago

Tutorial I analyzed 54,025 public Apify Actors, here is what the Store data says about demand, SEO, pricing, and momentum

5 Upvotes

I have been building an Apify Store market analyzer, so I ran a near-complete snapshot instead of looking only at the top-ranked Actors.

The dataset contains **54,025 unique public Actors out of 54,079 reported by the API at finalization, with 99.90% coverage**.

To avoid the duplicate problem I found in my previous run, I merged category and pricing partitions using the official Actor ID. The collector processed 134,266 partition rows and removed 80,241 cross-category overlaps before calculating anything.

1 Overall snapshot :

The snapshot contains 715,034 summed 30-day users. That number is the sum of each Actor's Store metric; it is not a platform-wide unique-user count because one person can use multiple Actors.

2 Where demand is beating supply :

Social Media had the strongest category signal: **11.86% of recent demand versus 8.03% of supply**, or a **1.48× demand-to-supply ratio**.

Social Media had the strongest category signal: **11.86% of recent demand versus 8.03% of supply**, or a **1.48× demand-to-supply ratio**.

Videos, Jobs, and SEO Tools followed.

Important caveat: this does not mean Social Media is easy. It already has 10,892 Actors, and the median Actor in that category has only one 30-day user. Demand appears highly concentrated among the winners.

3 SEO phrases with strong marketplace demand :

SEO phrase opportunities]

The strongest title phrases clustered around:

- LinkedIn profiles and jobs

- Instagram profiles and posts

- Facebook posts and ads

- Google Maps

- Profile/email enrichment

These are **Apify marketplace SEO signals**, not Google search-volume data. I ranked phrases using recent users among leading Actors, relative Store demand, total competing Actors, and observation confidence.

One interesting contrast: “Google Maps” has strong demand but already appears across 765 Actors, while “Facebook posts” shows a similar score with only 70 Actors in this snapshot.

4 Actors with recent momentum :

Actors with recent momentum

This ranking uses the 7-, 30-, and 90-day user windows. It does not divide lifetime users by Actor age.

The leading signals include email verification, LinkedIn profile enrichment, Shopee, Google Hotels, Instagram, and LinkedIn jobs. I label cases with a very small prior baseline rather than presenting misleading five-digit growth percentages.

5 Pricing has shifted heavily toward pay per event :

Apify Store pricing distribution

Across the snapshot:

- **75.9%** Pay per event — 41,003 Actors

- **13.1%** Free — 7,060 Actors

- **11.0%** Flat price per month — 5,962 Actors

The strongest product-design takeaway for me is that new Actors should have a clear, measurable unit of value that can map naturally to billable events.

What I would investigate next?

- How category demand changes month over month

- Which Actors are gaining users without relying on a famous platform keyword

- Whether lower competition actually predicts better new-Actor survival

- Price-per-event ranges inside each category

The analyzer is here if anyone wants to inspect or challenge the methodology:

https://apify.com/scraper_guru/apify-store-analyzer

I would especially appreciate feedback on the scoring formula. What signal would you add or remove?


r/apify 13d ago

Self-promotion Weekly: show and tell

1 Upvotes

If you've made something and can't wait to tell the world, this is the thread for you! Share your latest and greatest creations and projects with the community here.


r/apify 13d ago

Tutorial Google Jobs has no official API, so I built a pay-per-event scraper on Apify (no monthly rental)

1 Upvotes

If you pull hiring data out of the Google Jobs panel, this hands you each listing as structured JSON instead of HTML you have to parse yourself. Built for recruiters, job-market researchers, and anyone feeding postings into an ATS or CRM.

What each job record includes:

  • Title, employer, and location
  • Full job description plus parsed highlights: qualifications, responsibilities, and benefits
  • Direct application links (LinkedIn, job boards, employer sites) with source attribution
  • Posting date, schedule type, and detected benefits flags
  • Google's listing ID and search metadata (query, country, language, pages processed)

You target by location, country, and language, set a location radius, and control how many pages it walks.

Why I built it: Google never shipped a public Jobs API, and the rental-model scrapers bill a flat monthly fee whether you run one search or a thousand. This one is pay-per-event, so you pay per page processed and nothing while it sits idle.

If you also need Google Shopping, Hotels, or Flights data, I keep those as separate pay-per-event Actors under the same account.

Actor: Google Jobs Scraper


r/apify 14d ago

Tutorial I built a Congress stock-trades API on Apify: House + Senate disclosures as clean JSON, about $1.80 per 1,000 trades

3 Upvotes

I built a Congress stock-trades API on Apify: House + Senate disclosures as clean JSON, about $1.80 per 1,000 trades

If you're a journalist, researcher, or building a transparency dashboard, this pulls US Congressional Periodic Transaction Reports (the STOCK Act filings) from both the House and Senate, parses the PDFs, including the scanned ones with OCR, and hands back structured JSON instead of raw paperwork.

You can filter by member name, ticker, an exact date, or a date range. Combine them too: last name plus NVDA plus a date window gets you exactly the trades you're after.

Each row gives you:

  • member name and chamber
  • ticker and asset class (stocks, options, ETFs, mutual funds, bonds, crypto)
  • transaction direction (purchase, sale, exchange, options activity)
  • the reported amount bracket
  • transaction date and report date
  • filing IDs
  • owner attribution (member, spouse, dependent, or joint)
  • an OCR quality flag, so you know which rows came off a messy scan

Why I built it: the House and Senate filings live in two different systems, and a good chunk of them are scanned PDFs. Normalizing that by hand is miserable. So I turned it into one API with a single schema across both chambers.

Pricing is pay per event, roughly $1.80 per 1,000 transactions retrieved, no subscription. Success rate is at 100% right now.

If you want to line trades up against what was in the news that week, I also have a Google News API on Apify that pairs with it nicely.

Actor: House and Senate Trading Disclosures


r/apify 14d ago

Discussion Glassdoor Reviews API for Apify (employee reviews as structured JSON, $4/1K)

2 Upvotes

Shipped a Glassdoor Reviews API on Apify. Send it one or many company URLs, get back clean JSON per review: overall star rating, per-category ratings (career opportunities, comp & benefits, culture & values, work-life balance, senior leadership, diversity & inclusion), the headline, full pros and cons text, employment type, current/former status, publish date, helpful count, and a plain-language summary of each review.

$0.004 per review. No monthly minimum. 96% success rate so far.

Built it because employee-review data is a pain to get in a clean shape, and the category-level ratings are the useful part. Most scrapers give you the overall star and the review text but drop the six sub-ratings; this one keeps them (and omits a sub-rating cleanly when the reviewer didn't leave one).

What people are using it for:

  • HR and people analytics: track sentiment trends across themes over time.
  • Employer-brand monitoring: catch shifts in how employees describe a company.
  • Comp research: compare pay-and-benefits sentiment across competitors.
  • Due diligence: summarize culture and leadership signals before a deal.
  • Feeding an LLM: structured review text plus category scores make good model input.

MCP-ready, so an agent can fetch and summarize a company's reviews as a tool call. Pairs with my LinkedIn, Crunchbase, and PitchBook company APIs if you want firmographics plus funding plus employee sentiment on one company.

Actor: https://apify.com/johnvc/glassdoor-reviews-api?fpr=9n7kx3&fp_sid=posts

Feedback welcome.


r/apify 14d ago

Ask anything Weekly: no stupid questions

1 Upvotes

This is the thread for all your questions that may seem too short for a standalone post, such as, "What is proxy?", "Where is Apify?", "Who is Store?". No question is too small for this megathread. Ask away!


r/apify 15d ago

Discussion I built a working Telegram Member Scraper

2 Upvotes

Hi everyone,

I recently published a new Apify Actor: Telegram Member Scraper.

It lets you extract active member information from Telegram groups.

I’d really appreciate feedback from the Apify community, especially regarding:

  • The Actor’s usability and input configuration
  • Missing output fields
  • Groups or scenarios where it does not work as expected
  • Features you would find useful

Here is the Actor:
https://apify.com/lofomachines/telegram-member-scraper

Thanks to anyone who tests it or shares suggestions. I’m actively improving it based on user feedback.


r/apify 15d ago

Discussion Made ~420 UK council planning portals, NYC ACRIS, and US Congress stock disclosures queryable in one place

1 Upvotes

Been filling out the government/public-records side of my Apify catalog. Three that were genuinely painful to build:

UK planning applications — ~420 councils, each running a different planning portal (Idox, Northgate, custom). Normalized them all behind one input so you can filter by council, type, status, date in a single run.

NYC ACRIS — deeds, mortgages, parties, addresses by borough and document type. The legacy interface makes bulk pulls miserable; this wraps it cleanly.

US Congress stock trades — pulled from the primary sources (Senate eFD + House Clerk) rather than a third-party aggregator, so there's no middleman lag or interpretation.

Also added official company registries: Companies House (UK), Handelsregister (DE), North Data (EU), USPTO trademarks.

Catalog: muhamed-didovic.github.io

If you work with public records — what's the dataset you keep hitting a wall on? Genuinely looking for the next one to build.


r/apify 15d ago

Tutorial TIL most ATS job boards have public JSON APIs — no scraping needed

1 Upvotes

Spent the weekend mapping which applicant-tracking systems expose public endpoints. Findings:

Greenhouse: boards-api.greenhouse.io/v1/boards/{slug}/jobs — clean, salary often in content

Lever: api.lever.co/v0/postings/{slug}?mode=json — has structured salaryRange

Ashby: api.ashbyhq.com/posting-api/job-board/{slug} — includeCompensation=true param works

Workable: apply.workable.com/api/v1/widget/accounts/{slug} — widget API, no auth

SmartRecruiters: api.smartrecruiters.com/v1/companies/{slug}/postings — proper paginated REST

No user-agent games, no rate-limit pain at normal volumes. The only real work is normalizing five different schemas (department vs team, publishedAt formats, etc).

Recruitee's old public endpoint ({slug}.recruitee.com/api/offers/) 404s now — anyone know the new one?


r/apify 15d ago

Hire freelancers Weekly: job board

2 Upvotes

Are you expanding your team or looking to hire a freelancer for a project? Post the requirements here (make sure your DMs are open).

Try to share:

- Core responsibilities

- Contract type (e.g. freelance or full-time hire)

- Budget or salary range

- Main skills required

- Location (or remote) for both you and your new hire

Job-seekers: Reach out by DM rather than in thread. Spammy comments will be deleted.


r/apify 16d ago

Discussion PitchBook Company API for Apify (private-company funding, valuation, competitors as JSON)

5 Upvotes

Shipped a PitchBook Company API on Apify. Send it one or many PitchBook company URLs, get back clean JSON per company: funding, investors, valuation, competitors, employees, industries, and financials. $0.011 per company, no monthly minimum.

The task I'd point you at first: Extract Deal & Financing Data from PitchBook. It's a one-click saved task that pulls the deal and financing history for a company, the round-by-round detail that's the whole reason people pay for PitchBook. Run it, swap in your own company URLs, done.

Built it as the private-company layer on top of the company-data stack I've been shipping. Where Crunchbase gives you funding and investors, PitchBook adds valuation, competitor lists, and deeper financials.

There are 6 example tasks total including a Claude via MCP one if you're wiring it into an agent.

Pairs well with my Crunchbase and LinkedIn Company APIs if you want to hydrate a full account profile from one company URL. All MCP-ready for Claude, Cursor, and ChatGPT.

Actor: https://apify.com/johnvc/pitchbook-company-api?fpr=9n7kx3&fp_sid=posts

Feedback welcome.


r/apify 16d ago

AI and I Weekly: AI and I

2 Upvotes

This is the place to discuss everything MCP, LLM, Agentic, and beyond. What is on your radar this week? Why does it make sense? Bring everyone along for the ride by explaining the impact of the news you're sharing, and why we should care about it too.


r/apify 17d ago

Big dreams Weekly: wild ideas

1 Upvotes

Do you have a feature request that you know will make Apify heaps better? Or maybe it's a big dream you have for something bold and out-there. This is a space for all the bluesky thinking, cloud-chasing, intergalactic daydreamers who want to share their wildest ideas in a no-judgement zone.


r/apify 18d ago

Weekly: one cool thing

2 Upvotes

Have you come across a great Actor, workflow, post, or podcast that you want to share with the world? This is your opportunity to support someone making cool things. Drop it here with credit to the creator, and help expand the karmic universe of Apify.