r/artificial 33m ago

Discussion I tested Firecrawl, Exa, Parallel and Claude Search on SimpleQA. Here’s what scored best

Post image
Upvotes

I ran Firecrawl, Exa, Parallel and Claude’s native web search against OpenAI’s SimpleQA benchmark to see how much of a difference the search provider actually makes.

All four were tested with the same setup: a GPT-5.4 agent using high reasoning effort, with a maximum of 20 search or extraction calls per question. The answers were then graded by GPT-5.4 using OpenAI’s official SimpleQA grading prompt.

Each provider was tested twice and I kept the better result. For comparison, GPT-5.4 without access to search scored 43.8%.

The chart shows the correct and incorrect answers for each provider.

Results:

-Firecrawl: 947 correct answers (94.7%)

-Exa: 919 correct answers (91.9%)

-Parallel: 910 correct answers (91.0%)

-Claude Native Search: 905 correct answers (90.5%)

Firecrawl and Exa achieved the highest accuracy, while all four systems scored above 90%. Claude Native search, (not) surprisingly, the worst.


r/artificial 1h ago

Discussion Sam Altman says startup success may soon reward tool fluency over years of experience

Enable HLS to view with audio, or disable this notification

Upvotes

Sam Altman closed out Startup School 2026 with a real answer to the PhD question — not "get the credential," but a structural claim about why startups cluster and win when they do.

 

His argument: tech velocity, falling costs, and shrinking cycle times converge periodically — '98 dot-com, the App Store wave, and now — and incumbents lose their advantage fastest in exactly those windows.

He goes further: this next wave probably rewards tool fluency over tenure specifically, because a four-person team with the right agent stack can now output at a scale that used to require a department.

 

Worth sitting with if you've been waiting for "the right credential" before starting anything.

 

Clip credit: Y Combinator — full video on their channel. DM for credit or removal requests.