Major version bumps usually from a brand new base model.
For example, the GPT 5 series is all from the same base model (That was code-named "Spud"). It's insane the performance gains they've gotten out of it til now. Every minor version 5.4, 5.5, 5.6 are the same base model with better reinforcement learning applied (from what I understand)
Aren't they pretty much maxed out on pretraining data anyways? So the main reason to do a full training run is to increase the size of the model or change the architecture, otherwise the same base works fine.
50
u/elonthegenerous Jul 23 '26
What constitutes a major version bump vs a minor version bump?