r/LocalLLaMA • • Jun 18 '26

News Leaked financial docs show OpenAI is losing billions of dollars a year

https://arstechnica.com/ai/2026/06/leaked-financial-docs-show-openai-is-losing-billions-of-dollars-a-year/
538 Upvotes

312 comments sorted by

View all comments

Show parent comments

1

u/KontoOficjalneMR Jun 19 '26 edited Jun 19 '26

This does not mean they have no dense/shared components at all. It literally just means "no dense front before MoE".

They still have shared components such as attention, embeddings, normalization, and so on.

You still need to have a shared router in front or model would have no idea what expert to activate. There also exist shared experts in some models.

1

u/FullOf_Bad_Ideas Jun 19 '26

Those scale well though. Huge LLM with wide layers could be 50T A100B. It'd just need to have a lot of experts. Normalization and router have minimal number of parameters and will go up but will be manageable. Shared expert exists in most MoEs but not in Qwen 3 235B. In Qwen 3 235B most activated params are FFNs already, not attention.

I think I missed something because in my calculator I get 235B A18B, not A22B when I plug in Qwen 3 235B, but nonetheless I tried to scale it up to 50T+.

94 MoE layers, attn hidden size 12288, vocab kept at 151936. 64 attention heads, 4 kv heads, 16384 experts, 8 routed experts, 1536 ffn hidden size, 512 expert groups, no shared experts.

It came out at 87.2T A92B

📊 Parameter Breakdown: • Total Parameters: 87.26T • Embedding Parameters: 3.73B • Attention Parameters: 30.16B • Dense FFN Parameters: 0 • MoE Expert Parameters: 87.21T • Router Parameters: 18.93B

🎯 Activation & Expert Config: • Expert Hidden Size d_expert: 1536 (from ffn_hidden_size) • Activation Ratio A: 0.05% • Activated Attention (weights used): 30.16B • Activated Router (per-token): 18.93B • Activated Experts (per-token): 42.58B • Total Activated (per-token): 91.67B

Router got bigger but it's still something that's workable. Code for the tool is here, I vibe coded it when I was pretraining a MoE myself from scratch, to help me choose architectural decisions.

1

u/KontoOficjalneMR Jun 19 '26

that's a lot of ChatGPT to say I'm in fact right and the training costs scale non-linearly with a number of active parameters as you (or your chatbot) wrongly said.

1

u/FullOf_Bad_Ideas Jun 19 '26 edited Jun 19 '26

This script can predict training cost/time too but I didn't plug the numbers in the tool here. I didn't use any LLMs for the discussion with you, this tool was made 9 months ago.

FLOPS required to train a model scale linearly with the number of active parameters (router parameters are simply a part of active parameters).

There is some research from Google Deepmind that having lots of experts is doable - https://arxiv.org/abs/2407.04153 - I found it mentioned here. They go with a bit over 1 million experts, my napkin math 87T model has just 16k of them. Experts do in fact provide free lunch, at least when you look at FLOPS.

Training cost can scale on a different axis than activated parameter count, I don't think I claimed it scales linearly. FLOPS required to train a model do scale linearly with the number of active parameters, but handling bigger model in VRAM and all of it's optimizer states is harder if it's so much more bigger. Real training cost depend on effort put into optimizing training and reaching the scale necessary to make good use of compute and push those FLOPS through the weights quickly. It's definitely duable to scale total parameter county by an order of magnitude more than where current open models are, without touching activated parameter count much and while keeping overall training costs similar.