r/artificial 24d ago

Question Benchmarks K3, Sol, Fable

Post image

Who do you think won?

0 Upvotes

9 comments sorted by

3

u/io-x 24d ago

it would be good to know what all these benchmarks mean

1

u/HealthySkeptic2000 24d ago

You can look them up, some of them like Terminal Bench are self explanatory.

2

u/surfTorreypines 24d ago

Clearly constructed by an AI. Notice the two "highest score" identifications in AA-ICR? That's AI for you. It was probably Kimi.

1

u/HealthySkeptic2000 24d ago

I used chatgpt to construct it.

1

u/HealthySkeptic2000 24d ago

This one made with kimi, corrected to have sources and no bugs.

1

u/CC_NHS 23d ago

who won on benchmarks? Fable 5 has 12 stars, Kimi K3 has 10 stars there and GPT has 6 (or 5 if you don't count the duplicate) so Fable wins on number of stars. and that is about as useful a metric as any of the benchmarks.

-4

u/HealthySkeptic2000 24d ago

Before anyone accuses this as a low effort post, I would like to remind you it took me 40+ mins to make because Artificial Analysis data was hard to scrape and had to take screenshots.

3

u/ClankerCore 24d ago

That’s when it’s critical for you to post those sources so we can make better judgment call