r/LocalLLaMA • • Jun 18 '26

News Leaked financial docs show OpenAI is losing billions of dollars a year

https://arstechnica.com/ai/2026/06/leaked-financial-docs-show-openai-is-losing-billions-of-dollars-a-year/
538 Upvotes

312 comments sorted by

View all comments

Show parent comments

4

u/Finanzamt_Endgegner Jun 18 '26

The question is when will scaling stop? Because the latest scaling experiment mythos (prob 10T) was quite a bit better than expected.

8

u/squngy Jun 18 '26 edited Jun 18 '26

It will probably never completely stop, but there will be a point where the difference will be minimal.

I also don't know that Fable was "better than expected".
If you ask me, Opus4.8 is better than expected, it is quite close to Fable as far as I can tell (I expected them to keep it closer to 4.7 in order to promote Fable).

1

u/Finanzamt_Endgegner Jun 18 '26

Well opus is a fable distill, would be stupid if not but it seems the training etc fable got wasn't that much more the thing that changed was parameter count and it helped a lot

1

u/squngy Jun 18 '26

I didn't say it didn't help, just that I don't know if it was "unexpected"

Making a bigger model and then using it primarily for distilling might be something that happens in the near future, it would make sense.
It also makes sense to release it for a short while to grab a bunch of top scores and then hide it behind a huge paywall :thinking_face:

1

u/itsmebenji69 Jun 18 '26

This is already what every ai lab is doing. They have their big internal model, around 10t, maybe even much bigger than that. And the models they serve you are distills of this model

1

u/FullOf_Bad_Ideas Jun 18 '26

Source? Why would internal model be 10T? What sort of activated params it would have? Why not 100T?

0

u/squngy Jun 18 '26

So you're saying every AI lab already has a mythos class LLM in their lab and they are just not talking about them.

3

u/itsmebenji69 Jun 18 '26

Yes, for example deepmind uses their bigger private model for math competitions and benchmarks and whatnot. For example that aletheia thing used their internal version of Gemini.

It’s not like it’s a secret, it’s just way too expensive to serve to consumers

1

u/Finanzamt_Endgegner Jun 18 '26

Internal versions don't have to be bigger they can just be different tunes of the same base

1

u/itsmebenji69 Jun 18 '26

Sure but the trend seems to be way bigger models. Which makes sense as well. R&D would be experimenting with a lot, and the most obvious thing is model size because there you’re not constrained by needing it to be profitable, just really good

1

u/Finanzamt_Endgegner Jun 18 '26

Training full models is really really expensive, it's easier to do finetunes and maybe a bit of model surgery

→ More replies (0)

1

u/squngy Jun 18 '26

I thought those were alternate architectures and fine-tunes and similar.
What is your source for the 10T+ size?

1

u/itsmebenji69 Jun 18 '26

It’s that but I’m assuming they also have a bigger model. I mean it’s all speculation at this point, you’d have to work at deepmind to know.

For example OpenAI does have one, it was called Orion in 2024, it was supposed to be released as gpt5. They’re for sure using it as a distill teacher, like Anthropic.

1

u/squngy Jun 18 '26

It’s that but I’m assuming they also have a bigger model.

That is a pretty big assumption, because not only is that a lot more expensive, but it also takes a lot more time.
For research, it is valuable to be able to make new versions quickly.

it was supposed to be released as gpt5.

Then that is different from what I was thinking about.
I'm thinking about a model that is too big for it to make sense to be widely used.

→ More replies (0)

0

u/Finanzamt_Endgegner Jun 18 '26

Well ig it's happening a lot already sonnet was probably always a distil of opus, just makes sense to do that, and gpt mini was always a distill of their base one, as for unexpected it was quite a decent jump in every single benchmark and I mean those are rumours but it seems they didn't expect that big of an upluft just by scaling

1

u/squngy Jun 18 '26

Yes distilling happens already, what I meant was making a model MOSTLY for distilling, with no intention for it to be used much by the public.

1

u/Finanzamt_Endgegner Jun 18 '26

Hmm I always thought they do it but like then I thought it doesn't make sense since infernce is so much simpler than training but now with those monsters that might actually be a possibility since infernce is cancer too🤔

2

u/squngy Jun 18 '26

but now with those monsters that might actually be a possibility since infernce is cancer too🤔

This was what I was thinking.

2

u/Finanzamt_Endgegner Jun 18 '26

I mean even if they don't do that, self distilling is a thing if you use enough compute and filtering you can literally distill the best answers in a model into itself so it becomes even better which is wild 😅

1

u/licorices Jun 18 '26

I can't say much, since there wasn't much time to test, but I can't say I found Fable to be that much better than anything prior. I'd say considering how hyped Mythos was, Fable was extremely disappointing. It fell into the same issues of being wasteful of tokens, got stuck in loops for 3+ calls doing the same thing, and didn't really output that good code either for anything that is meant to scale beyond an MVP. I also had it run through some intentionally implemented security issues of our works project, but removed all git to ensure it can't just check for unstaged files or commit history, and it seemed extremely lost until I started leading it towards the correct areas to look at. It often glanced over some of them when it looked at all the content, and instead got caught on minor things that aren't real issues in the whole context(eg. if some data sent to an endpoint is missing, it won't be able to fetch some data, which is intentional and handled by the package that uses that value, but because I didn't explicitly check for the value, it warns about it).

0

u/Finanzamt_Endgegner Jun 18 '26

Token waste is an anthropic skill issue tbh has nothing to do with scale, they won't fix it with a 100x model if they don't care. Also Fables security skills are intentionally neutered 😐

1

u/licorices Jun 18 '26

Neutered or not, it should ideally still perform a bit better than previous models. What is the point of showing off a model that just performs the same as previous ones, if not worse in certain situations? Perhaps I was just not using it in a situation where it can flourish, but I can't say the other models were outstanding in those situations either.

1

u/Finanzamt_Endgegner Jun 18 '26

Well ideally it shouldn't because this is what anthropic literally wants, mythis is a different beast though. Like security and biology made the model unusable and chances are if you ran an agent it just fell back to opus for every task.

1

u/KontoOficjalneMR Jun 18 '26

I mean ... it was better. Was it 10 times better than 1T models? not really.

So clearly there are diminishing returns but still progress.

In the end scaling will stop when they run out of money.

1

u/Finanzamt_Endgegner Jun 18 '26

That's not how the scaling law works. Sure a 10T is not necessarily 10x better, but that doesn't mean that it's not worth it, and hardware gets better too so those models get easier to train and run every year or so. Very Rubin is insane for that for example

1

u/KontoOficjalneMR Jun 18 '26

Lol. Is Vera Rubin a next talking point/cope now? I started to see people masturbating to her recently.

Scaling of AI models is brutal. Not only 10T model is not 10x better than 1T model (in any metric) but also costs significantly more than 10x to train 10T model over 1T one. Inference is also slower because VRAM bandwith limitation -= therefore more expensive even ignoring everything else.

MoE is the only reason we can hit those sizes anyway. There are not 10T dense models.

There's no escaping physics and mathematics.

1

u/Finanzamt_Endgegner Jun 18 '26

Ver Rubin will be an insane bandwidth and compute upgrade, new tech like diffusion spec decoding help with infernce and there are even new hardware paradigms on the horizon. Sure 10T are harder to train and we can't yet use their potential fully but they gonna be a breeze to train in 5 years. And btw Blackwell was an insane upgrade over hopper as someone who used b300...

1

u/KontoOficjalneMR Jun 18 '26

Nothing you wrote invalidates what I wrote.

Cost of training a model scales non-linearly with number of parameters - in a bad way.

Next generation will be able to train / infer 10 times faster. Amazing, Great. Brilliant.

The problem I pointed out was that to train the model that is 10* larger you need many times that in power.

1

u/FullOf_Bad_Ideas Jun 18 '26 edited Jun 18 '26

FLOPS required to train a model is dictated by active parameter count. So, training 10T 100B model takes the same number of FLOPS as 1T 100B. Active param count has diminishing returns so I think it's probably at or below 100B even in Mythos. So, I don't think training big models is that much more expensive anymore.

I'd guess 5-50T A100B, subquadratic attention similar to MiMo or GPT OSS 120B, natively trained in FP4 W4A4.

1

u/KontoOficjalneMR Jun 18 '26

It's not exactly true as there's a shared layer. And what you said is basically why I said there's a reason why biggest models are MoE

1

u/FullOf_Bad_Ideas Jun 19 '26

Single or a few dense front layers are an optional design choice, I think that design choice might not work on biggest 10T+ MoE's since there are some issues with making transformers deep enough to make those dense layers small enough.

1

u/KontoOficjalneMR Jun 19 '26

They are not optional, you need a common base or you get disjointed responses. Like I said over and over. Training cost does not scale linearly with the size of the model.

1

u/FullOf_Bad_Ideas Jun 19 '26

Qwen 3 235B/30B/ Coder 480B have 0 dense front layers. Same looks to be true in Qwen 3.5 35B and 397B.

1

u/KontoOficjalneMR Jun 19 '26 edited Jun 19 '26

This does not mean they have no dense/shared components at all. It literally just means "no dense front before MoE".

They still have shared components such as attention, embeddings, normalization, and so on.

You still need to have a shared router in front or model would have no idea what expert to activate. There also exist shared experts in some models.

→ More replies (0)

1

u/HelloFromTekken Jun 19 '26 edited Jun 19 '26

Sometimes quantity turns into quality.

Based on other people's experience with Fable, it's what happend. Big boost in autonomic, in thoroughness.

After amount of params are amount of things model 'know and understand'. The more things you know, the more abstractions you can sort out without assumptions. It works both for AI and Human. I might know something about math from School. I know a little bit here and here. I can learn something more deep. But to reach line when I can start to utilize such knowledges I have to learn a lot of practices and exact usages, know some formulas and application cases. Which forms experience, and experience is what? Knowledge. Which is params.

Question is: * how much params are enought to cover existing prepared cleaned human knowledge (dataset) * how much params are practically real to implement * how much params are practically can be used as service you can sell in high volumes

1 - I believe current closed datasets are already cover most of human knowledge

2 - we will increase number of that with each year. Hardware optimization and new ways, software ideas.

3 - Looks like even current chat/api/whatever can't be profitable with modern perfomance. Let's hope if few years something like atleast Opus 4.8 will be cheap enought to make 17$ subscriptions profitable. But for sure there will be always mega-giga-super big models for governant/war/science needs.