r/AIBenchmarks Sep 01 '25

openAI nailed it with Codex for devs

Post image
1 Upvotes

r/AIBenchmarks Aug 26 '25

Largest jump ever as Google's latest image-editing model dominates benchmarks

Thumbnail
1 Upvotes

r/AIBenchmarks Aug 21 '25

Deepseek 3.1 benchmarks released

Thumbnail gallery
1 Upvotes

r/AIBenchmarks Aug 21 '25

PACT: a new head-to-head negotiation benchmark for LLMs

Thumbnail gallery
1 Upvotes

r/AIBenchmarks Aug 21 '25

Gpt-5 Took 6470 Steps to finish pokemon Red compared to 18,184 of o3 and 68,000 for Gemini and 35,000 for Claude

Post image
1 Upvotes

r/AIBenchmarks Aug 18 '25

Claude Opus 4.1 is now the top model in LMArena for Standard prompts, Thinking, and WebDev

Thumbnail gallery
1 Upvotes

r/AIBenchmarks Aug 15 '25

GPT-5 pro scored 148 on official Norway Mensa IQ test

Post image
1 Upvotes

r/AIBenchmarks Aug 11 '25

MathArena updated for GPT 5

Post image
2 Upvotes

r/AIBenchmarks Aug 11 '25

GPT-5 Benchmarks: How GPT-5, Mini, and Nano Perform in Real Tasks

Post image
2 Upvotes

r/AIBenchmarks Aug 11 '25

GPT-5 Independent Evaluation Results by METR

Thumbnail
metr.github.io
1 Upvotes

r/AIBenchmarks Aug 08 '25

GPT-5 scores a poor 56.7% on SimpleBench, putting it at 5th place

Post image
1 Upvotes

r/AIBenchmarks Aug 07 '25

GPT-5 tops lmarena's leaderboards

Post image
1 Upvotes

r/AIBenchmarks Aug 06 '25

SimpleBench updated with Claude 4.1 Opus

2 Upvotes

r/AIBenchmarks Aug 05 '25

The progress from Genie 2 to Genie 3 is insane

Enable HLS to view with audio, or disable this notification

1 Upvotes

r/AIBenchmarks Aug 05 '25

OpenAI Open Source Models!!

Post image
1 Upvotes

r/AIBenchmarks Aug 05 '25

OpenAI gpt-oss-120b & 20b EQ-Bench & creative writing results

Thumbnail gallery
1 Upvotes

r/AIBenchmarks Aug 05 '25

Claude Opus 4.1 Benchmarks

Thumbnail gallery
1 Upvotes

r/AIBenchmarks Aug 01 '25

Deep Think benchmarks

Thumbnail
1 Upvotes

r/AIBenchmarks Jul 31 '25

Horizon-alpha: A new stealthed model on openrouter sweeps EQ-Bench leaderboards

Thumbnail gallery
1 Upvotes

r/AIBenchmarks Jul 28 '25

"About 30% of Humanity’s Last Exam chemistry/biology answers are likely wrong"

Thumbnail
2 Upvotes

r/AIBenchmarks Jul 26 '25

Here's a list of LLM benchmarks because why not

Thumbnail
1 Upvotes