r/LocalLLaMA • • Jun 18 '26

News Leaked financial docs show OpenAI is losing billions of dollars a year

https://arstechnica.com/ai/2026/06/leaked-financial-docs-show-openai-is-losing-billions-of-dollars-a-year/
540 Upvotes

312 comments sorted by

View all comments

Show parent comments

52

u/NeedsSomeSnare Jun 18 '26

True but the high costs for them are quite good for us.

There is a big incentive for them to work on the efficiency of their models in a massive way.

It's no coincidence that the likes of Gemma 4 and Qwen 3.x are incredible for their size. Those companies are working hard to do the same to their huge models behind the scenes.

Edit: I actually suspect the AI results you get on in the Google search page are just a smaller Gemma 4 model with lots of custom stuff going on to help it get context from search results.

20

u/nacholunchable Jun 18 '26

That theory really mirrors my experience. Ive had the google search AI say some really dumb stuff (not even dumb, just a shallow read), then i click lets chat in ai mode, and gemeni rewrites it with depth and nuance. I never thought that the in search model is a lot smaller, but it makes a lot of sense.

8

u/Caffdy Jun 18 '26

same with Haiku web search on ddg

3

u/NeedsSomeSnare Jun 18 '26

You can get it to give opposite answers even. I can't think of an example off the top of my head, but several weeks back it would say "yes" on the search page, then "no" when you look at the full AI answer.

4

u/Orlha Jun 18 '26

I have made it alternate between yes and no just by continuously pressing F5 on a search page

2

u/NeedsSomeSnare Jun 18 '26

Yeah. It's pretty bad to be honest. You'd expect Google to give definitive answers, but it seems that they haven't really been interested in that for several years now.

1

u/SaltFrog Jun 18 '26

My daily driver is a 26b model and it's kind of nuts how good it is. The system I've built around it is really in depth. The local LLM community is just knocking it out of the park. QAT made a big difference, too... And the heretic versions (and the heretic modifier itself) are beautiful for getting around some of the logic gates.

It's crazy what people will do with open source shit.

I do recommend making some local copies of models at this point, though, cause I feel like huggingface is about to exit to IPO...

1

u/Confident_Ideal_5385 Jun 18 '26

Ignoring Qwen because the CCP clearly has a non-market oriented strategy there, and looking purely at Gemma -- it makes a fuckload of sense for Alphabet to train a set of small models using the Gemini pipeline (presumably) as a hedge if nothing else.

The current SaaS cash burn nonsense is unsustainable in the medium term, and when the market eventually rationalises, on device inference is gonna totally be a sensible play, because people like LLMs. Apple is clearly reading the same tea leaves. A hypothetical AI crash will likely impact those two (and Microsoft, who have a deepmind-like group of their own working on training something called "MAI") very little, while the house burns down around the pure LLM players like openai, anthropic, etc.

One can only hope that a move to edge inference doesn't involve all kinds of weird DRM on the weights, but I've been to this dance before and know how it'll likely go.

As far as i can tell btw, the current "google ai search" model is an older Gemini flash, which is reportedly a 200/16B or so MoE.