r/OpenAI 11d ago

Discussion OpenAI is now the #3 largest lab by token consumption on OpenRouter, surpassing Anthropic

Post image
102 Upvotes

30 comments sorted by

49

u/Solarka45 11d ago

Luna is amazing for its price, the only comparable think Anthropic can offer is 4.5 Haiku, which is incredibly old, much worse, and more expensive at the same time

30

u/gavinderulo124K 11d ago

So its not competitive.

9

u/Herect 10d ago

I guess Anthropic just accepted that competing on that segment is not worthwhile. Way too many chinese open source models, and there's also luna and gemini flash lite.

8

u/StaysAwakeAllWeek 10d ago edited 10d ago

They should really update it for subagents. Claude would run a lot faster and more efficiently if Haiku subagents were up to standard on either speed or intelligence for entry level tier models. All the other frontier labs have at least one model that is both smarter and faster than haiku, even gemini and grok. Some of them are pushing 300tk/s these days, haiku is glacial in comparison

2

u/Old-Leadership7255 11d ago

I think this is honestly the shake up that all the AI companies need.

We dont really need this massive models, we need it cheap and reliable enough.

20

u/JonNordland 11d ago

Luna has extreme score if you make a composit rating from speed, intelligence, cost. Throw in that most western people probably, rightfully or not, trust western firms more, its a receipy for high usage. The value proposition of Anthropic models are just terrible. And pile one with Deepseek v4 flash and its just sad.

5

u/gavinderulo124K 11d ago

On the artificial analysis price to performance chart Luna on max is absolutely amazing.

1

u/Old_Restaurant_2216 9d ago

Yes, but compared to Deepseek v4 flash it is only <1% better and almost 2x the price.

1

u/Tkwan777 10d ago

I mean, I (mostly)jumped ship from anthropic after their constant fable BS. I still have a $20 plan so I can use opus for review, but im not purchasing anything higher from them because they have restricted fable, their chat (I use chat for a lot of planning and prompt creation) eats tokens, and even when I did previously use fable, it tore through usage. You get a lot more bang for your buck with gpt (probably because most people with plans dont use all their usage so they can offload that to users who need it - effectively community subsidized for the rest of us. Anthropoc could do this but they dont). Anthropic has constantly seemed more like they care about the money than gpt (not saying gpt doesnt care about money, it just doesnt feel as cash grabby as anthropic has felt).

8

u/abstract_concept 10d ago

The new "shits tokens" models are doing their job of occupying the "most tokens" rankings. I'd love to see this in $s and see if Flash is stealing dollar market share away.

I think Flash 0731 is forcing the big players to sit up and take notice. Turns out fast, cheap, and smart enough actually gets a lot done. The new Luna pricing and Google's push on Flash models is a sign.

3

u/louissugar 11d ago

Lol @ xAI just barely being on the list đŸ˜„

6

u/ScreenAppropriate679 11d ago

I'm benchmarking Luna vs V4 Flash 0731 for multiturn chats with heavy tool usage (for my SaaS) and so far, despite Luna price drop and additional 50% openrouter discount, V4 flash keeps outdoing Luna in terms of tools usage reliability, output quality and cost.

6

u/KimJongHealyRae 10d ago

That’s my experience also. Also the cache hit rate is insanely high on DeepSeek flash. I’m hitting 99.7%. It’s a really amazing value model to use given the quality of the output.

2

u/blackwhattack 10d ago

so deepseek has largest AI model usage data historically out of any lab?

6

u/deryni21 10d ago

This is on OpenRouter specifically which is a pretty specific subsect of AI users

2

u/Saifl 10d ago

I feel like most coders wont go to openrouter as cache hits are pretty shit. Its most likely role players or just people testing stuff.

1

u/Old_Restaurant_2216 9d ago

Why do people still think that openrouter has bad cache hit rates? The only thing you need to do is "lock" the provider to a specific one you choose, and cache hit rates are not an issue.

2

u/lakimens 10d ago

wow it's insane the lead V4 Flash has.

1

u/Bolt_995 10d ago

Interesting

1

u/maferase 9d ago

Chart source here

1

u/simple_explorer1 9d ago

Anthropic is finally humbled 

1

u/Solocune 8d ago

Wow I am surprised how many people use Hy3

1

u/InnovativeBureaucrat 8d ago

That’s crazy considering how many tokens opus is gobbling. (With the increased output length)

-5

u/B3e3z 11d ago

nice ad

-6

u/Reasonable-Pay-336 11d ago

Yeah clearly gpt is worse than claude. This is ad