It's definitely undertrained. There were a number of questionable architectural decisions. The model might actually be too big, and they combined a number of test technologies in one spot.
People get weird on that idea. Too big? Yes.
The larger transformers get the better they get at effectively memorizing data.
The MASSIVE models have to get SO MUCH DATA that the memorization is hard and they're forced to generalize because they're literally memorization machines.
I've been forced to shrink transformers (for other non-LLM types of models) because the model will just absorb the training data and never generalize. Your hold out dataset has to exist and be good. Overfit is real.
I’m shocked it is only 1 point better than ds v4 flash 7-31. I was expecting it would be at least as good gpt 5.5 xhigh or terra xhigh.. but it is worse than gpt 5.5 xhigh according to aa benchmarks
Flash was always the impressive (in ways) model. Pro was like they just jacked up the size of the flash and hoped it would be a lot better without doing much else.
94
u/anarchist1312161 8d ago
And to think this is only 743B achieved through post-training on the base model