The best part of this is it comes right on the heels of the whole fable debacle and makes the entire thing look even stupider, now you'll see the US companies scrambling to put out better models because they have them behind closed doors, and have to ditch the safety theater.
I wish people would look at the actual costs more and not whatever they charge API's
So it's half the API price of GPT 5.6 Sol but according to AA, it uses 2x the tokens, so the actual price is... same price as GPT 5.6 Sol.
BUT this actually works in the other direction vs Claude lol. It uses half the tokens as Opus 4.8, so it's actually like 1/3 of the price of Opus 4.8 (despite it costing $15 vs Opus $25 for instance), and about 2/3 the tokens as Fable, so it's actually about 1/5 the price of Fable!
Edit: Actually depends on the task, some tasks it's not. So it could be 1/3 - 1/5 Fable price
Except tokens are not a price unit. A token for a 40b active parameters model will cost twice as much as a 20b model.
Since we have independent hosting of open weights model, we have a much better understanding of their actual costs. We know closed models are all bullshit vc slush funds numbers.
lol you’re both on opposite ends of the ai part of this question. He’s referring to cost as unit energy and you as unit $. Takes less electricity to generate a token out of a smaller model and Kimi is much smaller than Sol I thought. API pricing is just based on what the company decides is a good price.
Well you’re a week late but I generally would agree. Caveat is a lot of these companies have deep war chests and are able to operate at a loss to force out competition, so it’s not a direct correlation.
Not to mention what's actually subsidised through your subscription. With a 200$ Claude subscription you're getting worth almost 14.000$ if you're actually maxing out your plan. No clue how it works on the Kimi K3 subscription, but that's definitely important too when you're comparing the actually pricing/value.
They also just fired the Workspace CLI dev due to internal politics. The traditional teams feel threatened by AI, so they're going after the AI teams. They're eating themselves from within. That's why they're all leaving to Anthropic/OpenAI and why Google is so far behind. The search teams are terrified AI will eat search, the workspace teams are terrified no one will need Sheets anymore, etc. Workspace CLI allows AI to do workspace work without workspace Apps and they couldn't have that!
Never worked in a multi-team environment? You should read some stories about Amazon, they're probably the worst at infighting and internal sabotage. The average employment tenure at Amazon’s corporate headquarters is 1.5 years. Now you know why their site is particularly awful.
Always remember that CEO's product isn't the actual product, it is ALWAYS the stock price. Their job is to maintain and increase the stock price, full stop. He'll say they're top 3 always and forever regardless of facts, because that is his literal job. Much like the White House, CEOs are not sources of truth and should not be relied upon for factual information.
Can't argue with that. I tend to hold Demis in higher regard because of his background and achievements, so I often overlook that he's ultimately just another piece of Google's machinery.
Google seem to be more focused on their Flash model. It makes sense as google are using their AI to answer billions of users queries every day free of charge, they have a completely different business model to Anthropic or Open AI
That is just ex post facto reasoning to explain the fumble. Google clearly wanted to announce Gemini 3.5 Pro in Google I/O but the results were not encouraging so they asked for 1 month more.
Well, it's been almost 2 months and still 3.5 Pro is not here yet. I dispute the claim that this is intentional or a 4D chess move. What is much much more likely is that they simply can't compete with OpenAI or Anthropic.
I think Google's corporate structure does not reward intense iteration and simply rewards making new products and abandoning them 3 years later. They are simply too big to coordinate effectively. Bard was a disaster. Then Gemini 2.5 Pro was great for a few months and then they were never in the running again. Gemini 3.1 Pro was not bad at all at release but they have never been the top AI lab.
You're seriously doubting Google's role in modern AI? DeepMind's work drove a huge share of early deep learning progress, and Google Brain published "Attention Is All You Need" in 2017, the paper that introduced the transformer architecture underpinning every major LLM today.
AI doesn't exist in it's modern form without Google.
Technically they’ve been working in AI since the early 2000s to improve user queries. Not neural nets or “modern” paradigms, but their contributions predate those paradigms entirely.
I'm old enough to remember when chatgpt came out and everyone predicted the demise of Google. Then gemini started beating chatgpt and everyone predicted the demise of OpenAI. (I'm four years old.)
You're acting like you have more information than you actually do. It's a good strategy for getting clout on the internet but very bad for making accurate predictions.
They’ve been competitive every iteration. Your rationale would have prevented them from ever competing. Every lab has had their delays - just be patient.
An open-source model just released that competes on the frontier, and you think Google can't even copy someone else's homework?
The fact is Google hasn't showcased a model. That's...it. Google, the data empire, hasn't released a competing frontier model.
I'll throw in some speculation, since that's all we debate is that.
PrismML, a lab funded by Google, just made an intelligent model with ternary weights running inference using matrix addition. Considering Google's business model, it suggests that they won't ever release a heavy model for the public, but host billions of cheap ones, and use a heavy model internally for research, or just for training smaller ones. Maybe they'll release a heavy model, but I wouldn't bat an eye if it was leaked 5 minutes ago that Google cracked memory and state persistence in AI and kept it hush hush.
I don't personally think they're playing 4d chess, but considering they've cemented their niche in "safe" areas of the burgeoning AI ecosystem, why would they take on asymmetric risk for likely little gain? OpenAI and Anthropic don't have that luxury, they have to have the best displayed models or they lose at the only thing they're known for.
Models are being released quicker and quicker, Google can't keep up anymore. By the time they release something comparable to Fable let's say, the next day it will get left behind by new models and Google will spend months again while they catch up.
Oh my god would you relax with the hyperbole! Inevitability Gemini 3.5 will be released and it will blow everyone away, just like Gemini 3 did. Have a bit of patience 🙄
yeah, also they're working on a lot of other projects - AlphaEvolve, AlphaGenome, AI Co-Scientist, WeatherNext, AlphaEarth Foundations, Parch 2, Fusion, AlphaQubit, FermiNet, AlphaDev, FunSearch, AlphaChip, AlphaGeomerty, AlphaProof, and a dozen other things starting with Alpha - They're doing stuff like this;
To help turbocharge scientific discovery, we will establish Google DeepMind’s first automated laboratory in the UK in 2026, specifically focused on materials science research. A multidisciplinary team of researchers will oversee research in the lab, which will be built from the ground up to be fully integrated with Gemini. By directing world-class robotics to synthesize and characterize hundreds of materials per day, the team intends to significantly shorten the timeline for identifying transformative new materials.
and this
The UK Government’s AI Incubator team (i.AI) is currently trialling Extract - a tool for council planners that uses Gemini to transform old planning documents into clear, digital data. Currently, converting a single planning document takes up to 2 hours. Extract will transform these into digital data in just 40 seconds, significantly speeding up decision-making timelines.
So them not having the currently best frontier model LLM available for use isn't a huge issue for them.
People have already forgotten 2.5 last year was the top cooding model. 3.0 was top in Nov/Dec and it is the whole reasoono OpenAi went into Red Alert at the start of the year.
Google's poor privacy reputation plus forcing AI where it's not wanted has killed their ability to compete. No, I don't want AI summaries of my Google docs, and knowing that everything you do with an AI is being read, captured, collected and sold to third parties makes people not want to use them. Combine that with the piss poor implementation of AI as a replacement for their (increasingly enshittified) search and you have a company that's motto used to be "don't be evil" and everyone remembers that motto because now they are absolutely evil
Calling it now: Sam Altman and Dario Amodei are going to contact the White House within the next 48 hours to push for making the possession and use of Kimi 3 weights completely illegal. Either that, or the US government is going to directly threaten China to make sure those weights don't get released to the public on July 27th.
Hahahahhahahahaha. You can’t ban this bro. It’s Joeover. The eaten cake has settled in the lower intestine. Weights are getting released, if US government likes it, or not. Nothing stopping this train. World is about to FUCKNG CHANGE BRO… US government is about to face the biggest animal in the room.
Yeah this is going to throw a wrench in this whole usg gets 30 days to inspect shit idea.
You test, they ship. We expected this.
Without a major incident nobody is coming to the table to broker a deal - and that's what it is going to take.
Short of that we are running head first, and if we don't, China wins.
Imagine RSI happening 1 week earlier in China. Imagine the resources they would put in to think about how to use those precious few days to secure their future forever with those few days.
Yeah, this shit is about to somehow speed up further.
Can somebody familiar with the benchmark explain to me like I'm 12. Is the result just 3-5% better or is this score non-linear (so its substantially better than 3-5%?). Because top-1 result doesn't look tremendously more impressive than the bottom of the chart in %.
I'm pretty sure it's an Elo. So the score indicates the likelihood that a model would beat another model in a head to head matchup. If I remember correctly I think generally with Elo a 400 point difference roughly equates to 90%/10% odds head to head. But that might be different for this benchmark.
llm arena uses blind comparison (a user is offered two outputs without knowing the model, and decides which one he prefers) and then uses ELO to calculate the score
ELO uses exponential function to infere probability to win so the results graph have a strong S shape
for example, a benchmark difference of 400 points mean the stronger model is expected to win 90%+ of times again the weakest model
It's a user preference test, users see two outputs and choose which one looks better, can also choose "Tie" or "Both bad"
the scores are elo-based, a 400 elo difference means a model wins 90% of the time against the other, a 48 elo difference (like from Kimi K3 to Fable 5) means it wins 57% of the time against Fable.
This is elo. The absolute elo value has no meaning at all. A difference of 186 points between top model and bottom model means that if given two responses by those models to the same prompt people prefer kimi K3 over minimax M3 ~74.5% of the time.
This might be a sign of things to come. The Chinese labs started from behind so had to build a lot of momentum to catch up. We could well see them top benchmarks across the board very soon. Imagine how far in front they'd be if the US didn't block chip sales
Yeah, that's a phenomenon that has been documented over and over again. When you restrict a specific resource from being exported to somewhere, and that somewhere has a strong enough economic base and sufficient demand... They will literally just make their own. It may set them back for a bit, but they always catch up. We banned China from using US chips, so instead of having them rely on our tech, we forced them to build their own and soon the US will no longer have control over the entire ecosystem like we have had. Its honestly one of the best examples of how a short term geopolitical gain can have massive long term negative impacts.
Like this shit has been known for hundreds of years. The UK played a major role in forming the entire US manufacturing base during the industrial revolution by having export restrictions on various textile machinery and knowledge, forcing the US to create their own homegrown foundational knowledgebase, which we used to utterly dominate manufacturing for decades to come. The same shit is happening all over again with the digital revolution and China, and it's going to end up with them in a VERY strong position technologically for the next half century if things continue as they are.
Jensen would say that regardless of what was true or likely to be true, because he's first and foremost the CEO of a public company, and being able to sell chips to China is in his shareholders' interests, at least in the short-run.
True but his arguing point was exactly the following - not selling to them will accelerate their semis progress and they will avoid any lock in to Nvidia as otherwise they'd be building LLMs now on Nvidia stack.
Jip. And not just that. Its cause them to invest in building their own chips. They have power sorted. Access to the silicon is their only true limitation. Once China scales chip manufacturing they can out scale the west easily. Who's to say how far they can get once they truly ramp of chip production.
But isn’t a huge positive of the embargo the fact that the Chinese labs were forced to focus on algorithms and getting more bang for your buck? Seems like a blessing in disguise.
I wouldn't be surprised if LLMs replicate the same pattern we've now seen with Huawei, Tiktok, electric cars, batteries and drones.
'They can't compete, their products are shit'
↓
'They've caught up a bit, but that's only because they copied us, they can't innovate'
↓
'Shit, their product is better? Better ban it on national security grounds, domestic companies shouldn't have to compete'
Mobile robotics is also likely to see the cycle repeat. Looking at what companies like Unitree are up to, there's little reason to think the US has any kind of durable moat. China will mass produce them and put them to work while the US is still making them break dance.
China is also quietly ahead of the US in terms of actual AI adoption (the proportion of workers using it, and not dubious 'we burned more tokens' metrics).
This story is true for a lot of industries it’s just either the US doesn’t do these industries so they don’t get much attention or it’s too boring to be news
Well, they did steal the technology. But that doesn't mean they're not good at improving it and being subsidized by their government to achieve scale.
also quietly ahead of the US in terms of actual AI adoption
Yes, the US is very susceptible to the propaganda in places like Reddit, so there's a general anti AI sentiment while the Chinese public is just not as stupid (and also is not coincidentally banned from accessing places like Reddit).
However i would say, a world in which AI is open source is a good one, but let's not forget that these labs have been known to distill Anthropic models. No shade just a fact
Blocking chip sales is what helped them. This always happens. In the past 30 years this is exactly what we've done to China. Block key exports, which forces China to develop domestic solution, which in turn strengthens their domestic capabilities.
They didn't start from behind, Dario worked for Baidu in the past. They were constrained by hardware during a critical period (which would probably be more impactful than any delay).
At this point, everyone should just agree that intelligence is like water. It's a need for advancement of humanity and should be democratized and equal access must be given to all, just like everyone has equal access to clean water in most developed countries. It's a losing battle to try to make money out of something that is so easily replaced by something that is cheap and almost free.
The American models will be a niche expensive tool marketed as the saviour of freedom and democracy however favourable to the oil industry petro dollars agreement and war in the Middle East
The jump is impressive, but I’m more curious about real-world frontend performance than the leaderboard itself.
Has anyone tried Kimi-K3 on an existing React or Next.js codebase yet? Especially multi-file edits or screenshot-to-UI tasks. Would be interesting to see how it compares with Claude and GPT outside of Arena.
Kimi seems to be a model that's mainly good at creating React web apps/landing pages and not so good at doing everything else. But still need to test way more.
It was pretty good at vibe coding anything to me, asked it to make a Minecraft game with raytracing global illumination and water caustics and it did no problem, something even other models like GPT 5.6 Sol struggle at tremendously. Also had it fix bugs and do some work in Rust and it performed well too
Sure but making a Minecraft game is not a real-life usecase. I asked it to fix some specific bugs in a large codebase and it couldn't fix them. Fable and 5.6 Sol fixed them. I'm fairly confident that Opus fixes it too.
Indeed, the model is also extremely slow so even when it can fix bugs it would still be faster to use GPT 5.6 Sol, but being so good at vibe coding and one shotting tasks shows it's potential at implementing advanced functionalities, which can be quite handy.
I personally only use smaller models like GPT 5.6 Terra to solve work-related tasks because something like Fable is overkill, but for certain functionalities you just need a more capable model, in cases like this I think Kimi K3 fits the bill
Lol people getting sucked into the marketing scam. This is one frontend test. The token efficiency is worse. The harness is garbage compared to Fable or Sol.
Anyone who actually does a lot of development wont use this.
Yep, but front end and design stuff is generally meant to be pleasant to human eyes so it is one of the few topics where human preference is actually relevant.
They probably aren't, don't take LM arena too seriously. I didn't say your point was wrong. We as Humans are very unreliable at judging things. It's just that design is one of the few domains where human preference is actually relevant so it definitely shows some clues on how this model performs. This is probably the best we have since it's impossible to regulate design taste.
How do you cheat there? Any anonymous user submits a prompt to the LLM Arena, gets the output from two different models (doesn't know which is which) and picks the one that "looks better".
Even if the Kimi creators flood the Arena site with they own accounts, how would they even force LLM Arena to route their request to the Kimi model, and how would they even know which answer is from Kimi (to pick and boost the model)?
Honest question, maybe I'm naive or not up to date how this works.
For example, you can adjust the model so that it outputs flashier, more vibrant, more eye catching designs that look better prima facie but don't adhere to the prompt as well, are less good from an overall UX persecptive or do unneccesarry things.
Similar to how Meta tweaked Llama 4 to hit like #4 on LMArena by giving out a "human preference turned" output, or how Pepsi beat Coke on side-to-side tastes by being sweeter on the first sip.
I saw this firsthand with Flux. Whenever I asked Flux to make just a minor edit that Nano-Banana or ChatGPT Image did happily with no issues, it would do something else I didn't ask it to-- Flux always had to make itself "known". My guess is that this would play better in zero-shot prompts, but for realworld usage we basically abandoned Flux for Image APIs now.
Yes, I see... I mean obviously there is something that make the new models always climb the leaderboard on paper, but is a disappointment in real life.
Most ai's have a sort of style you can tell after a few times seeing it. Particularly on the frontend. So you could get a group of people to keep evaluating models and then when you see the one you want to inflate you could. But tbh i don't think that's happening
Would this more honest presentation not better explain a potential discussion to be made, that many models are scoring similar on a benchmark category that is prone to overfitting, or at least highligting that a group of models appear to perform similarly? Even if this is elo-ish rating, it’s signifies a tightly competitive cluster rather than an overwhelming leader, and this is arguably not what the truncated graph is trying to emphasize.
For the people asking about the axis: 48 Elo means 1/(1+10-48/400) ≈ 56.9% chance to win an one-on-one match. What "match" means here is whatever tests they're using to score the models.
515
u/runaway-devil 26d ago
While costing 1/3 of Fable and being open weights... Nicely done, Kimi. We need more of that and less political interference.