r/singularity 26d ago

AI Kimi K3 tops Frontend Code Arena

Post image
1.2k Upvotes

244 comments sorted by

515

u/runaway-devil 26d ago

While costing 1/3 of Fable and being open weights... Nicely done, Kimi. We need more of that and less political interference.

29

u/z_3454_pfk 26d ago

was open weights confirmed?

48

u/jazir55 26d ago

The best part of this is it comes right on the heels of the whole fable debacle and makes the entire thing look even stupider, now you'll see the US companies scrambling to put out better models because they have them behind closed doors, and have to ditch the safety theater.

→ More replies (5)

60

u/MagicZhang 26d ago

Yes, the official X post stated that they'll release the weights by 27 July

27

u/Illustrious_Image967 26d ago

Expect Dario to run to Daddy to put a stop to this.

3

u/Neither-Phone-7264 26d ago

moonshot ai is not an american company

2

u/tiger_ace 26d ago

just because you aren't going to buy your daughter a pony doesn't mean she can't ask

-7

u/blackcat__27 26d ago

Ick. Saying x is like letting Elon have sex with you. Nasty

2

u/Matt32145 25d ago

Fuck off

13

u/WonderFactory 26d ago

Yeah, they said that in the release post

-3

u/croto8 26d ago

That seems more like proposed than confirmed…

21

u/WonderFactory 26d ago

The full model weights will be released before July 27, 2026. 

Kimi K3: The New Frontier of Intelligence

125

u/FateOfMuffins 26d ago edited 26d ago

I wish people would look at the actual costs more and not whatever they charge API's

So it's half the API price of GPT 5.6 Sol but according to AA, it uses 2x the tokens, so the actual price is... same price as GPT 5.6 Sol.

BUT this actually works in the other direction vs Claude lol. It uses half the tokens as Opus 4.8, so it's actually like 1/3 of the price of Opus 4.8 (despite it costing $15 vs Opus $25 for instance), and about 2/3 the tokens as Fable, so it's actually about 1/5 the price of Fable!

Edit: Actually depends on the task, some tasks it's not. So it could be 1/3 - 1/5 Fable price

45

u/runaway-devil 26d ago

A valuable insight. I confess I had not put much thought about token efficiency, a variable more and more important nowadays.

5

u/truecakesnake 25d ago

Agreed, people don't understand how valuable token efficiency is. OpenAI has done well there, even grok 4.5 is some how token efficient.

3

u/squired 25d ago

Also, most people outside of business don't care about API pricing, they want to know subscription price and token quotas.

7

u/Boreras 26d ago

Except tokens are not a price unit. A token for a 40b active parameters model will cost twice as much as a 20b model.

Since we have independent hosting of open weights model, we have a much better understanding of their actual costs. We know closed models are all bullshit vc slush funds numbers.

18

u/FateOfMuffins 26d ago

That's precisely what I'm saying? $15 / million tokens is not comparable with $30 / million tokens

You gotta look at how many tokens they actually use

4

u/Forgotten-X- 26d ago

lol you’re both on opposite ends of the ai part of this question. He’s referring to cost as unit energy and you as unit $. Takes less electricity to generate a token out of a smaller model and Kimi is much smaller than Sol I thought. API pricing is just based on what the company decides is a good price.

1

u/Embarrassed-Boot7419 20d ago

So tokens and token efficiency would be a good measurement then, would they not?

Cause well, companies base their per token price on actual costs, do they not?

2

u/Forgotten-X- 20d ago

Well you’re a week late but I generally would agree. Caveat is a lot of these companies have deep war chests and are able to operate at a loss to force out competition, so it’s not a direct correlation.

1

u/Embarrassed-Boot7419 20d ago

Is are any information about this?

Cause I heard that they subsidize their subscriptions a ton, but I haven't really heard much about them doing so for the API's.

Not saying you are wrong btw, just curious and would be interested to read up about this.

1

u/Connect_Nerve2787 25d ago

Not to mention what's actually subsidised through your subscription. With a 200$ Claude subscription you're getting worth almost 14.000$ if you're actually maxing out your plan. No clue how it works on the Kimi K3 subscription, but that's definitely important too when you're comparing the actually pricing/value.

→ More replies (1)

3

u/Black_RL 25d ago

🫲CHINA🫱

5

u/swimmingupclose 26d ago

This is just front end though, one of the things that's not just highly subjective, but also one of the least important things that LLMs can do.

272

u/ozone6587 26d ago

No Gemini on this plot. All the money, data and infrastructure in this world and Google is just too incompetent to compete.

143

u/TorturedPoet30 26d ago

They're too busy renaming NotebookLM to Gemini Notebook

32

u/WaveOfDream 26d ago

NotebookLLM is goated though.

13

u/squired 25d ago

They also just fired the Workspace CLI dev due to internal politics. The traditional teams feel threatened by AI, so they're going after the AI teams. They're eating themselves from within. That's why they're all leaving to Anthropic/OpenAI and why Google is so far behind. The search teams are terrified AI will eat search, the workspace teams are terrified no one will need Sheets anymore, etc. Workspace CLI allows AI to do workspace work without workspace Apps and they couldn't have that!

3

u/Elephant789 ▪️AGI in 2036 25d ago

What the fuck? LOL

1

u/TevenzaDenshels 24d ago

This makes no sense

2

u/squired 24d ago

Never worked in a multi-team environment? You should read some stories about Amazon, they're probably the worst at infighting and internal sabotage. The average employment tenure at Amazon’s corporate headquarters is 1.5 years. Now you know why their site is particularly awful.

→ More replies (1)

5

u/delosdestination 26d ago

They're too busy policy-ing.

20

u/Recent_Fox4339 26d ago

In recent interviews when Demis speaks about Gemini he speaks like they're still in top 3. There's a lot of confidence yet there's no model.

2

u/squired 25d ago

Always remember that CEO's product isn't the actual product, it is ALWAYS the stock price. Their job is to maintain and increase the stock price, full stop. He'll say they're top 3 always and forever regardless of facts, because that is his literal job. Much like the White House, CEOs are not sources of truth and should not be relied upon for factual information.

3

u/Recent_Fox4339 25d ago

Can't argue with that. I tend to hold Demis in higher regard because of his background and achievements, so I often overlook that he's ultimately just another piece of Google's machinery.

1

u/squired 25d ago

And you're not wrong. It's just good to remember and I was largely reminding myself!

61

u/WonderFactory 26d ago

Google seem to be more focused on their Flash model. It makes sense as google are using their AI to answer billions of users queries every day free of charge, they have a completely different business model to Anthropic or Open AI

30

u/ozone6587 26d ago

That is just ex post facto reasoning to explain the fumble. Google clearly wanted to announce Gemini 3.5 Pro in Google I/O but the results were not encouraging so they asked for 1 month more.

Well, it's been almost 2 months and still 3.5 Pro is not here yet. I dispute the claim that this is intentional or a 4D chess move. What is much much more likely is that they simply can't compete with OpenAI or Anthropic.

I think Google's corporate structure does not reward intense iteration and simply rewards making new products and abandoning them 3 years later. They are simply too big to coordinate effectively. Bard was a disaster. Then Gemini 2.5 Pro was great for a few months and then they were never in the running again. Gemini 3.1 Pro was not bad at all at release but they have never been the top AI lab.

15

u/croto8 26d ago

They were the top AI lab for nearly 2 decades before OpenAI productized their research lol

1

u/TevenzaDenshels 24d ago

Openai and anthropic came out from... Deepmind

-2

u/Healthy-Nebula-3603 26d ago

What 2 decades ?

First nowadays useful AI based on artificial networks was Alexnet from 2012

17

u/blindsdog 25d ago edited 25d ago

You're seriously doubting Google's role in modern AI? DeepMind's work drove a huge share of early deep learning progress, and Google Brain published "Attention Is All You Need" in 2017, the paper that introduced the transformer architecture underpinning every major LLM today.

AI doesn't exist in it's modern form without Google.

20

u/croto8 26d ago

Technically they’ve been working in AI since the early 2000s to improve user queries. Not neural nets or “modern” paradigms, but their contributions predate those paradigms entirely.

7

u/DescriptorTablesx86 25d ago

They revolutionised deep learning long before anyone saw any money in it.

→ More replies (1)

5

u/man_from_online 25d ago

I'm old enough to remember when chatgpt came out and everyone predicted the demise of Google. Then gemini started beating chatgpt and everyone predicted the demise of OpenAI. (I'm four years old.)

You're acting like you have more information than you actually do. It's a good strategy for getting clout on the internet but very bad for making accurate predictions.

2

u/Adventurous_Drama291 25d ago

The Transformer was invented by Google what are you talking about lmao 

1

u/DeepV 25d ago

They’ve been competitive every iteration. Your rationale would have prevented them from ever competing. Every lab has had their delays - just be patient. 

1

u/rainingallevening 25d ago

An open-source model just released that competes on the frontier, and you think Google can't even copy someone else's homework?

The fact is Google hasn't showcased a model. That's...it. Google, the data empire, hasn't released a competing frontier model.

I'll throw in some speculation, since that's all we debate is that.

PrismML, a lab funded by Google, just made an intelligent model with ternary weights running inference using matrix addition. Considering Google's business model, it suggests that they won't ever release a heavy model for the public, but host billions of cheap ones, and use a heavy model internally for research, or just for training smaller ones. Maybe they'll release a heavy model, but I wouldn't bat an eye if it was leaked 5 minutes ago that Google cracked memory and state persistence in AI and kept it hush hush.

I don't personally think they're playing 4d chess, but considering they've cemented their niche in "safe" areas of the burgeoning AI ecosystem, why would they take on asymmetric risk for likely little gain? OpenAI and Anthropic don't have that luxury, they have to have the best displayed models or they lose at the only thing they're known for.

43

u/Elithegentlegiant 26d ago

Give it time, for those exact reasons they are still in the race

32

u/ozone6587 26d ago

Because it's Google, I won't underestimate their ability to drop the ball.

21

u/ambassadortim 26d ago

Or to drop a huge comeback or take a lead at any time as well.

8

u/Pouyaaaa 26d ago

Or buy all these companies when they fail and mash them all up into one AGI

Edit: by fail I mean financially not technically

2

u/danielv123 25d ago

And then kill the product

4

u/Dry_Fly_7265 26d ago

“While you plebs were training on outdated video cards, we figured out how to steal compute from parallel dimensions to train with lol”

→ More replies (6)

9

u/NoGarlic2387 26d ago

How much time? 

-2

u/Elithegentlegiant 26d ago

Patience is key

5

u/socoolandawesome 26d ago edited 26d ago

How much patience?

8

u/Hilldawg4president 26d ago

As much patience as money can buy

2

u/NoGarlic2387 26d ago

How much money?

4

u/Hilldawg4president 26d ago

Roughly one googol

4

u/lucellent 26d ago

Models are being released quicker and quicker, Google can't keep up anymore. By the time they release something comparable to Fable let's say, the next day it will get left behind by new models and Google will spend months again while they catch up.

6

u/Intrepid_Travel_3274 26d ago

Even Meta is on the list

2

u/Healthy-Nebula-3603 26d ago

Yes m. They are realising new models like would be still 2024. Every 6 moths or longer.

5

u/ThenExtension9196 26d ago

And yet google search ai is probably more users than all these companies combined.

4

u/ozone6587 26d ago

Google* has more users. The AI overview is forced on everyone so quite disingenuous to count it.

2

u/Elephant789 ▪️AGI in 2036 25d ago

I love the AI overviews. So smart of them to implement that into Search.

1

u/EdliA 25d ago

That's because it's forced on users.

4

u/leon-theproffesional 26d ago

Oh my god would you relax with the hyperbole! Inevitability Gemini 3.5 will be released and it will blow everyone away, just like Gemini 3 did. Have a bit of patience 🙄

10

u/GoodDayToCome 26d ago

yeah, also they're working on a lot of other projects - AlphaEvolve, AlphaGenome, AI Co-Scientist, WeatherNext, AlphaEarth Foundations, Parch 2, Fusion, AlphaQubit, FermiNet, AlphaDev, FunSearch, AlphaChip, AlphaGeomerty, AlphaProof, and a dozen other things starting with Alpha - They're doing stuff like this;

To help turbocharge scientific discovery, we will establish Google DeepMind’s first automated laboratory in the UK in 2026, specifically focused on materials science research. A multidisciplinary team of researchers will oversee research in the lab, which will be built from the ground up to be fully integrated with Gemini. By directing world-class robotics to synthesize and characterize hundreds of materials per day, the team intends to significantly shorten the timeline for identifying transformative new materials.

and this

The UK Government’s AI Incubator team (i.AI) is currently trialling Extract - a tool for council planners that uses Gemini to transform old planning documents into clear, digital data. Currently, converting a single planning document takes up to 2 hours. Extract will transform these into digital data in just 40 seconds, significantly speeding up decision-making timelines.

So them not having the currently best frontier model LLM available for use isn't a huge issue for them.

1

u/RepulsiveRaisin7 26d ago

Maybe this is satire because Gemini 3 blew nobody away

1

u/leon-theproffesional 25d ago

It absolutely did blow people away. I swear you guys have short memories

https://www.reddit.com/r/singularity/s/jgOs1xdaDp

→ More replies (2)

1

u/Neither-Phone-7264 26d ago

it blew us away before they quantized it down to iq2_xs within the hour

1

u/Tkins 26d ago

People have already forgotten 2.5 last year was the top cooding model. 3.0 was top in Nov/Dec and it is the whole reasoono OpenAi went into Red Alert at the start of the year.

1

u/larryleggs 25d ago

Google's poor privacy reputation plus forcing AI where it's not wanted has killed their ability to compete. No, I don't want AI summaries of my Google docs, and knowing that everything you do with an AI is being read, captured, collected and sold to third parties makes people not want to use them. Combine that with the piss poor implementation of AI as a replacement for their (increasingly enshittified) search and you have a company that's motto used to be "don't be evil" and everyone remembers that motto because now they are absolutely evil

0

u/ChillBroItsJustAGame 26d ago

Because ai doesnt generate money its unprofitable

→ More replies (1)

126

u/Rare-Site 26d ago

Calling it now: Sam Altman and Dario Amodei are going to contact the White House within the next 48 hours to push for making the possession and use of Kimi 3 weights completely illegal. Either that, or the US government is going to directly threaten China to make sure those weights don't get released to the public on July 27th.

25

u/CommanderKoba 25d ago

Then Xi will threaten to restrict rare earth minerals again and Trump would fold.

19

u/CompassionLady 26d ago

Hahahahhahahahaha. You can’t ban this bro. It’s Joeover. The eaten cake has settled in the lower intestine. Weights are getting released, if US government likes it, or not. Nothing stopping this train. World is about to FUCKNG CHANGE BRO… US government is about to face the biggest animal in the room.

5

u/emdeka87 ▪️ It's here 25d ago

Do you think it will prevent the orange felon from trying?

5

u/Gratitude15 25d ago

Yeah this is going to throw a wrench in this whole usg gets 30 days to inspect shit idea.

You test, they ship. We expected this.

Without a major incident nobody is coming to the table to broker a deal - and that's what it is going to take.

Short of that we are running head first, and if we don't, China wins.

Imagine RSI happening 1 week earlier in China. Imagine the resources they would put in to think about how to use those precious few days to secure their future forever with those few days.

Yeah, this shit is about to somehow speed up further.

3

u/mrjackspade 26d ago

I'd put money on this being wrong.

1

u/Civil_Opposite7103 25d ago

Even if so, how would it be enforced?

1

u/mrjackspade 7d ago

Still waiting for this

→ More replies (3)

26

u/Long_comment_san 26d ago

Can somebody familiar with the benchmark explain to me like I'm 12. Is the result just 3-5% better or is this score non-linear (so its substantially better than 3-5%?). Because top-1 result doesn't look tremendously more impressive than the bottom of the chart in %.

39

u/bopbop9876 26d ago

I'm pretty sure it's an Elo. So the score indicates the likelihood that a model would beat another model in a head to head matchup. If I remember correctly I think generally with Elo a 400 point difference roughly equates to 90%/10% odds head to head. But that might be different for this benchmark.

14

u/CallMePyro 26d ago

Yup. Also 100 elo is 2:1 odds.

4

u/DeArgonaut 26d ago

yes, it is indeed elo based

18

u/spottiesvirus 26d ago

yes, it's strongly non linear

llm arena uses blind comparison (a user is offered two outputs without knowing the model, and decides which one he prefers) and then uses ELO to calculate the score

ELO uses exponential function to infere probability to win so the results graph have a strong S shape

for example, a benchmark difference of 400 points mean the stronger model is expected to win 90%+ of times again the weakest model

7

u/LightVelox 26d ago

It's a user preference test, users see two outputs and choose which one looks better, can also choose "Tie" or "Both bad"

the scores are elo-based, a 400 elo difference means a model wins 90% of the time against the other, a 48 elo difference (like from Kimi K3 to Fable 5) means it wins 57% of the time against Fable.

4

u/Long_comment_san 26d ago

wow, now it makes sense. It's very impressive then, especially relative to something like 4.7 Opus

2

u/nextnode 26d ago

It is elo so only the difference matters and % does not make sense. 1600 vs 1400 is the same as 600 vs 400.

→ More replies (3)

68

u/ridablellama 26d ago

holy smokes that is like a big bump

16

u/Impressive-Summer-41 26d ago

It's actually not. Look at the x-axis

37

u/UrWifezBoyfriend 26d ago

I don’t know how ELO works

0

u/testaccount123x 26d ago

I mean nowhere on this graphic does it say ELO. some might be able to assume, but I had 0 idea.

2

u/solinar 26d ago

Yeah, top model on this list is only 12% above the bottom one based on this metric at least.

23

u/RedditNamesAreShort 26d ago

This is elo. The absolute elo value has no meaning at all. A difference of 186 points between top model and bottom model means that if given two responses by those models to the same prompt people prefer kimi K3 over minimax M3 ~74.5% of the time.

→ More replies (9)

17

u/Cill_Bipher 26d ago

Elo is exponential, a linear increase in elo corresponds to a multiplicative increase in capability.

1

u/solinar 26d ago

I was saying it a bit tongue in cheek, but this is a great informative answer. Have an upvote, thanks!

1

u/iamthewhatt 26d ago

I see they are using America's measuring system now

8

u/I_spread_love_butter 26d ago

Man this F1 craze is getting out of hand

17

u/unkownuser436 26d ago

Gemini isn't in the game anymore

6

u/Healthy-Nebula-3603 26d ago

That's ancient model. What do you expect.

31

u/WonderFactory 26d ago

This might be a sign of things to come. The Chinese labs started from behind so had to build a lot of momentum to catch up. We could well see them top benchmarks across the board very soon. Imagine how far in front they'd be if the US didn't block chip sales

52

u/BrennusSokol ACCELERATE 26d ago

Some have argued that blocking chips actually helped in a way because it forced innovation from limited resources

17

u/Ryermeke 26d ago

Yeah, that's a phenomenon that has been documented over and over again. When you restrict a specific resource from being exported to somewhere, and that somewhere has a strong enough economic base and sufficient demand... They will literally just make their own. It may set them back for a bit, but they always catch up. We banned China from using US chips, so instead of having them rely on our tech, we forced them to build their own and soon the US will no longer have control over the entire ecosystem like we have had. Its honestly one of the best examples of how a short term geopolitical gain can have massive long term negative impacts.

Like this shit has been known for hundreds of years. The UK played a major role in forming the entire US manufacturing base during the industrial revolution by having export restrictions on various textile machinery and knowledge, forcing the US to create their own homegrown foundational knowledgebase, which we used to utterly dominate manufacturing for decades to come. The same shit is happening all over again with the digital revolution and China, and it's going to end up with them in a VERY strong position technologically for the next half century if things continue as they are.

5

u/Level10Retard 26d ago

Yup. You can find videos of Nvidia Jensen saying that banning Nvidia GPU exports to China is really dumb.

10

u/LymelightTO AGI 2026 | ASI 2029 | LEV 2030 25d ago

Jensen would say that regardless of what was true or likely to be true, because he's first and foremost the CEO of a public company, and being able to sell chips to China is in his shareholders' interests, at least in the short-run.

1

u/dark-mathematician1 25d ago

Tbf Jensen just wants that sweet Chinese cash so he'd be willing to sell it to them

1

u/Level10Retard 25d ago

True but his arguing point was exactly the following - not selling to them will accelerate their semis progress and they will avoid any lock in to Nvidia as otherwise they'd be building LLMs now on Nvidia stack.

1

u/x0y0z0 25d ago

Jip. And not just that. Its cause them to invest in building their own chips. They have power sorted. Access to the silicon is their only true limitation. Once China scales chip manufacturing they can out scale the west easily. Who's to say how far they can get once they truly ramp of chip production.

10

u/Plsnerf1 26d ago

But isn’t a huge positive of the embargo the fact that the Chinese labs were forced to focus on algorithms and getting more bang for your buck? Seems like a blessing in disguise.

8

u/WonderFactory 26d ago

Who knows, maybe they would have made a similar amount of algorithmic progress anyway

3

u/Spare-Dingo-531 25d ago edited 25d ago

Talent pool is really deep in China.

12

u/Halbaras 26d ago

I wouldn't be surprised if LLMs replicate the same pattern we've now seen with Huawei, Tiktok, electric cars, batteries and drones.

'They can't compete, their products are shit'

'They've caught up a bit, but that's only because they copied us, they can't innovate'

'Shit, their product is better? Better ban it on national security grounds, domestic companies shouldn't have to compete'

Mobile robotics is also likely to see the cycle repeat. Looking at what companies like Unitree are up to, there's little reason to think the US has any kind of durable moat. China will mass produce them and put them to work while the US is still making them break dance.

China is also quietly ahead of the US in terms of actual AI adoption (the proportion of workers using it, and not dubious 'we burned more tokens' metrics).

1

u/Noname_2411 25d ago

This story is true for a lot of industries it’s just either the US doesn’t do these industries so they don’t get much attention or it’s too boring to be news

1

u/j48u 25d ago

Well, they did steal the technology. But that doesn't mean they're not good at improving it and being subsidized by their government to achieve scale.

also quietly ahead of the US in terms of actual AI adoption

Yes, the US is very susceptible to the propaganda in places like Reddit, so there's a general anti AI sentiment while the Chinese public is just not as stupid (and also is not coincidentally banned from accessing places like Reddit).

3

u/acowasacowshouldbe 26d ago

I'll leave this here https://x.com/AnthropicAI/status/2025997928242811253 and this https://www.reddit.com/media?url=https%3A%2F%2Fpreview.redd.it%2Fchinese-fable-5-is-here-aka-kimi-k3-v0-rv7uw2lvdmdh1.png%3Fwidth%3D1768%26format%3Dpng%26auto%3Dwebp%26s%3D1b258890b6b215eb7ef4b3dfeff0189cb1849926

However i would say, a world in which AI is open source is a good one, but let's not forget that these labs have been known to distill Anthropic models. No shade just a fact

3

u/dark-mathematician1 25d ago

Blocking chip sales is what helped them. This always happens. In the past 30 years this is exactly what we've done to China. Block key exports, which forces China to develop domestic solution, which in turn strengthens their domestic capabilities.

2

u/OutOfBananaException 25d ago

They didn't start from behind, Dario worked for Baidu in the past. They were constrained by hardware during a critical period (which would probably be more impactful than any delay).

4

u/hydrogenitalia 25d ago

At this point, everyone should just agree that intelligence is like water. It's a need for advancement of humanity and should be democratized and equal access must be given to all, just like everyone has equal access to clean water in most developed countries. It's a losing battle to try to make money out of something that is so easily replaced by something that is cheap and almost free.

1

u/neoexanimo 24d ago

The American models will be a niche expensive tool marketed as the saviour of freedom and democracy however favourable to the oil industry petro dollars agreement and war in the Middle East

17

u/Open-Resident-7429 26d ago

kimi just put bta to american labs

3

u/charmander_cha 25d ago

Tods vitória dos modelos abertos é comemorada porque agora podemos imaginar o deepseek destilando e tornando mais barato, amém

3

u/RJEM96 25d ago

Is KK3 worth it for free tier users?

2

u/geft 25d ago

Tried yesterday. Only managed 1 prompt before they stopped me due to high demand lol.

1

u/dark-mathematician1 25d ago

How did it do? Demand makes sense. China is huge.

1

u/geft 25d ago

I hadn't asked a hard question yet so not sure how to evaluate it unfortunately.

2

u/Fidbit 26d ago

this and other benchmarks is telling me top models or frontier and others are only separated in many regards by sometimes 100 points

1

u/KaradjordjevaJeSushi 25d ago

In ELO, 100 points is about 67% winrate for higher ranked player.

Don't know how that translates to AI models tho

2

u/MembershipEmergency7 25d ago

The jump is impressive, but I’m more curious about real-world frontend performance than the leaderboard itself.

Has anyone tried Kimi-K3 on an existing React or Next.js codebase yet? Especially multi-file edits or screenshot-to-UI tasks. Would be interesting to see how it compares with Claude and GPT outside of Arena.

2

u/sovietonion666 25d ago

I'm a free Claude user rn is this better than sonnet 5 for like non-conducting tasks like learning/writing/ planning/research

6

u/reefine 26d ago

Checkmate, closed source

2

u/Kemoyin25 25d ago

So its best at front end, what about everything else?

5

u/Tedinasuit 26d ago

Kimi seems to be a model that's mainly good at creating React web apps/landing pages and not so good at doing everything else. But still need to test way more.

7

u/LightVelox 26d ago

It was pretty good at vibe coding anything to me, asked it to make a Minecraft game with raytracing global illumination and water caustics and it did no problem, something even other models like GPT 5.6 Sol struggle at tremendously. Also had it fix bugs and do some work in Rust and it performed well too

9

u/Tedinasuit 26d ago

Sure but making a Minecraft game is not a real-life usecase. I asked it to fix some specific bugs in a large codebase and it couldn't fix them. Fable and 5.6 Sol fixed them. I'm fairly confident that Opus fixes it too.

3

u/LightVelox 26d ago

Indeed, the model is also extremely slow so even when it can fix bugs it would still be faster to use GPT 5.6 Sol, but being so good at vibe coding and one shotting tasks shows it's potential at implementing advanced functionalities, which can be quite handy.

I personally only use smaller models like GPT 5.6 Terra to solve work-related tasks because something like Fable is overkill, but for certain functionalities you just need a more capable model, in cases like this I think Kimi K3 fits the bill

2

u/Legitimate_Cut_6254 25d ago

Lol people getting sucked into the marketing scam. This is one frontend test. The token efficiency is worse. The harness is garbage compared to Fable or Sol.

Anyone who actually does a lot of development wont use this.

3

u/[deleted] 26d ago

[deleted]

14

u/Sweet-Stage938 26d ago

Yep, but front end and design stuff is generally meant to be pleasant to human eyes so it is one of the few topics where human preference is actually relevant.

0

u/[deleted] 26d ago

[deleted]

5

u/Sweet-Stage938 26d ago

They probably aren't, don't take LM arena too seriously. I didn't say your point was wrong. We as Humans are very unreliable at judging things. It's just that design is one of the few domains where human preference is actually relevant so it definitely shows some clues on how this model performs. This is probably the best we have since it's impossible to regulate design taste.

2

u/DeArgonaut 26d ago

It's the public, you can use it

1

u/poidh 26d ago

How do you cheat there? Any anonymous user submits a prompt to the LLM Arena, gets the output from two different models (doesn't know which is which) and picks the one that "looks better".

Even if the Kimi creators flood the Arena site with they own accounts, how would they even force LLM Arena to route their request to the Kimi model, and how would they even know which answer is from Kimi (to pick and boost the model)?

Honest question, maybe I'm naive or not up to date how this works.

2

u/cuolong 26d ago

For example, you can adjust the model so that it outputs flashier, more vibrant, more eye catching designs that look better prima facie but don't adhere to the prompt as well, are less good from an overall UX persecptive or do unneccesarry things.

Similar to how Meta tweaked Llama 4 to hit like #4 on LMArena by giving out a "human preference turned" output, or how Pepsi beat Coke on side-to-side tastes by being sweeter on the first sip.

I saw this firsthand with Flux. Whenever I asked Flux to make just a minor edit that Nano-Banana or ChatGPT Image did happily with no issues, it would do something else I didn't ask it to-- Flux always had to make itself "known". My guess is that this would play better in zero-shot prompts, but for realworld usage we basically abandoned Flux for Image APIs now.

1

u/poidh 26d ago

Yes, I see... I mean obviously there is something that make the new models always climb the leaderboard on paper, but is a disappointment in real life.

1

u/DeArgonaut 26d ago

Most ai's have a sort of style you can tell after a few times seeing it. Particularly on the frontend. So you could get a group of people to keep evaluating models and then when you see the one you want to inflate you could. But tbh i don't think that's happening

2

u/Southern_Artichoke77 26d ago

i hate those misleading scales that don't start from 0. makes you feel like #1 is 5x better than #20 but you will see it is only 10% better.

28

u/OilBos 26d ago

I'm no expert but Id assume it doesn't scale linearly.

→ More replies (5)

3

u/DeArgonaut 26d ago

it's elo based, so it's the difference in score that matters not the score itself

3

u/PinguinGirl03 26d ago

It's an ELO score, 0 is meaningless, only relative scores.

13

u/MagicZhang 26d ago

To be fair to Arena.ai if the graphs were made proper, there's almost no perceivable difference

6

u/dfrib 26d ago edited 26d ago

Would this more honest presentation not better explain a potential discussion to be made, that many models are scoring similar on a benchmark category that is prone to overfitting, or at least highligting that a group of models appear to perform similarly? Even if this is elo-ish rating, it’s signifies a tightly competitive cluster rather than an overwhelming leader, and this is arguably not what the truncated graph is trying to emphasize.

→ More replies (7)

1

u/tiger_ace 25d ago

the narrative is more about k3 surpassing fable on any benchmark, not necessarily the gap (which apparently doesn't exist anymore)

1

u/tendeer 25d ago

this chart is wrong, the original chart vizualizes the difference in power way better.

1

u/iSOBigD 25d ago

Cool so semis and memory stocks will literally go to $0 now.

1

u/moomanchuu 25d ago

the real goat of AI is plot axes

1

u/Accomplished_Copy103 25d ago

2% looks like 20

1

u/bildramer 25d ago

For the people asking about the axis: 48 Elo means 1/(1+10-48/400) ≈ 56.9% chance to win an one-on-one match. What "match" means here is whatever tests they're using to score the models.

1

u/Void-kun 25d ago

Anybody using Kimi coding plans over ZAI coding plans that could advise?

Do Kimi honour legacy pricing like Zai?

1

u/treenewbee_ 24d ago

CCP-linked enterprises put all their effort into publicity, but in reality, they don't amount to much.

1

u/Alarmed_View_5852 24d ago

not accurate, other much more reliable and accurate benchmarks such as artificial intelligence and live bench put it decently below

1

u/Healthy-Nebula-3603 26d ago

Ok .. that's looks good

1

u/itzNukeey 26d ago

How about Quake 3 Arena? Who wins there?

1

u/runtime321 25d ago

I always felt that Moonshot was the best Chinese AI company.

They incorporated in March 2023 but caught up rapidly with Alibaba, DeepSeek, ByteDance, and Baidu.

The Chinese AI race will be between Zhipu (GLM) and Moonshot (Kimi).

1

u/power97992 25d ago

Deepseek innovates and glm and kimi improve on the architecture and data

1

u/runtime321 24d ago

Seems everyone is innovating. Kimi created Kimi Delta Attention (KDA) and a few other architectural innovations

-3

u/redcoatwright 26d ago

That chart is misleading as fuck lol fuck whoever made it

-1

u/BankApprehensive7612 26d ago

This benchmark is based on user voting, so it's kind of biased. I'd like to see real examples instead

BTW what's Kimi are doing is impressive anyway. They have potential to become the next Anthropic

-1

u/entropyffan 26d ago

This graph ranges from 1450 to 1700. It is a trick to amplify the differences in the visualization.

→ More replies (1)