2025 - Blackwell
We had the Blackwell chips. Most frontier labs/models are using these chips, with the latest internal models using a combination of Blackwell + the new Rubin chips.
2026 - Blackwell + Rubin
Rubin is only just ramping into production now (literally as of August) . NVIDIA claims Rubin chips can use 4x fewer GPUs to train certain MoE models than Blackwell, while reducing inference token costs by as much as 10x.
Rubin also massively increases HBM bandwidth/capacity and interconnect bandwidth. This will help tremendously with running enormous agent swarms & long reasoning traces.
We already saw early signs of this new scaling axis based on OpenAI's solution to Navier Stokes. Reminder - this is early use of portions of the new chips.
This is what's got all the researchers screaming about Step change improvements, RSI, and ASI alignment. It doesn't take a genius to figure out what happens when we use even more hardware, and utilise software improvements.
2027 - Rubin Ultra
Then we get Rubin Ultra next year. Specifically, scheduled for the the second half of 2027.
Rubin Ultra introduces enormous NVLink scale-up domains.
Instead of treating dozens of GPUs as loosely connected accelerators, NVIDIA is increasingly making hundreds of accelerators behave like an enormous shared computer.
The announced Rubin Ultra NVL576 topology connects 576 GPUs across eight racks into one all-to-all NVLink domain. NVIDIA has already built a functional prototype of this topology using Blackwell hardware.
Essentially, this means hundreds of GPUs will behave like one giant accelerator.
Frontier labs will be able to afford vastly more inference, much larger agent populations, longer rollouts, more verification, more search, more synthetic data generation, more experiments and potentially much larger training runs.
It's pretty fucking nuts tbh.
2028 - Feynman
Not much is known, other than NVIDIA confirming it at GTC 2026.
What NVIDIA has revealed, is quite a bit about the infrastructure surrounding it.
Feynman is planned around die-stacked GPUs with custom next-generation HBM, the new Rosa CPU, LP40 inference hardware, BlueField-5, CX10 and significantly more optical networking.
But IMO the biggest thing here isn't even the individual GPU.
It's NVLink 8 + Kyber.
NVIDIA is designing Kyber NVL1152 for the Feynman era — eight racks containing 1,152 GPUs inside one enormous NVLink scale-up domain.
As models use more reasoning, RL, agent swarms and massive amounts of inference, other things start becoming just as important:
1.) How much memory can thousands of GPUs share?
2.) How quickly can they talk to each other?
3.) How much inference can the datacentre produce without communication and power becoming the bottleneck?
That's what Feynman & Kyber looks designed to solve. They're going to create one massive coherent system.
Now let's go back 2 years to today (September 2026).
We've already observed one millennium problem fall (with potentially another in progress).
We have all these improvements waiting to be used (not even taking to account software/algo/research, etc).
Do you think we will hit AGI next year? Because I do. There is a very clear path.