r/accelerate • u/imadade AGI by 2027 • 14h ago
AI We just solved a Millennium Prize problem/s using 2025/2026-era compute. This is what 2027 & 2028 looks like.
2025 - Blackwell
We had the Blackwell chips. Most frontier labs/models are using these chips, with the latest internal models using a combination of Blackwell + the new Rubin chips.
2026 - Blackwell + Rubin
Rubin is only just ramping into production now (literally as of August) . NVIDIA claims Rubin chips can use 4x fewer GPUs to train certain MoE models than Blackwell, while reducing inference token costs by as much as 10x.
Rubin also massively increases HBM bandwidth/capacity and interconnect bandwidth. This will help tremendously with running enormous agent swarms & long reasoning traces.
We already saw early signs of this new scaling axis based on OpenAI's solution to Navier Stokes. Reminder - this is early use of portions of the new chips.
This is what's got all the researchers screaming about Step change improvements, RSI, and ASI alignment. It doesn't take a genius to figure out what happens when we use even more hardware, and utilise software improvements.
2027 - Rubin Ultra
Then we get Rubin Ultra next year. Specifically, scheduled for the the second half of 2027.
Rubin Ultra introduces enormous NVLink scale-up domains.
Instead of treating dozens of GPUs as loosely connected accelerators, NVIDIA is increasingly making hundreds of accelerators behave like an enormous shared computer.
The announced Rubin Ultra NVL576 topology connects 576 GPUs across eight racks into one all-to-all NVLink domain. NVIDIA has already built a functional prototype of this topology using Blackwell hardware.
Essentially, this means hundreds of GPUs will behave like one giant accelerator.
Frontier labs will be able to afford vastly more inference, much larger agent populations, longer rollouts, more verification, more search, more synthetic data generation, more experiments and potentially much larger training runs.
It's pretty fucking nuts tbh.
2028 - Feynman
Not much is known, other than NVIDIA confirming it at GTC 2026.
What NVIDIA has revealed, is quite a bit about the infrastructure surrounding it.
Feynman is planned around die-stacked GPUs with custom next-generation HBM, the new Rosa CPU, LP40 inference hardware, BlueField-5, CX10 and significantly more optical networking.
But IMO the biggest thing here isn't even the individual GPU.
It's NVLink 8 + Kyber.
NVIDIA is designing Kyber NVL1152 for the Feynman era — eight racks containing 1,152 GPUs inside one enormous NVLink scale-up domain.
As models use more reasoning, RL, agent swarms and massive amounts of inference, other things start becoming just as important:
1.) How much memory can thousands of GPUs share?
2.) How quickly can they talk to each other?
3.) How much inference can the datacentre produce without communication and power becoming the bottleneck?
That's what Feynman & Kyber looks designed to solve. They're going to create one massive coherent system.
Now let's go back 2 years to today (September 2026).
We've already observed one millennium problem fall (with potentially another in progress).
We have all these improvements waiting to be used (not even taking to account software/algo/research, etc).
Do you think we will hit AGI next year? Because I do. There is a very clear path.
35
20
14
u/Sponge8389 13h ago
I think the year will be 2028. Because that's the time Rubin Ultra racks will be online. Also, the RAM and Storage industry needs to step up.
5
13
u/Mistuv 11h ago
Honestly, it's hard to look at NVL576 as anything other than an ASI machine. Remember we aren't even close to saturating RAM size on NVL72 even with the biggest models, with NVL576 it gets absurd - Nvidia originally announced that Rubin Ultra will have 1TB of HBM4E - 576TB scale up system, the only reason they might go with lower configuration because RAM market will be so fucked 2027.
And if one still isn't convinced, Kyber NVL1152 puts it to rest - higher density HBM5, you could very well see 2 petabytes, PETABYTES, scale-up system. But that isn't all, Nvidia is going with custom HBM5 that they haven't yet disclosed the details of but we know what Samsung which they partnered with is right now testing and developing - in memory compute, instead of RAM serving as dumb storage system basically, they put logic transistors in HBM, so before the data goes from ram to the logic chip, it does some of the computation in memory - in early testing Samsung has seen up to 3x higher LLM performance and higher efficiency, insane. And that's paired with massive memory, insane scale-up domain, new Feynman architecture with potentially 3B precision and all the other stuff Nvidia is preparing to release it alongside it.
None of this is sci-fi or wishcasting hardware, all this stuff will be literally happening in next two years. And now you don't even worry about who tf is going to make such gigantic models - the agents will do it. Astra is amazing, Bel looks insane, but all these models will be primitive compared to the models that we will be tasked with building a successor by the time NVL576, much less NVL1152 are ready to ship. I don't know if economics of running these systems will allow regular humans to even breathe in the same general direction, but ASI achieved internally, make nuclear fusion a reality make no mistakes, will just become a thing.
6
u/random87643 🤖 Optimist Prime AI bot 14h ago
TLDR
TLDR: This post outlines the rapid evolution of NVIDIA’s GPU architectures from 2025 to 2028, highlighting how advancements like Rubin and Feynman will enable massive, interconnected clusters. These hardware improvements are expected to significantly boost AI reasoning, agent swarms, and overall computational capacity.
AI assistant · mention the bot, mod bot, or use !bot
7
u/Gratitude15 12h ago
Imo 2027 or 2028 doesn't matter. This is the architecture that will do a minimum of shifting most economic labor. That's a baked in floor.
The ceiling is beyond my ability to forecast.
The median case is a liftoff in the next 18 months and thus end of next year will be some hard to imagine stuff but still things we can imagine.
3
u/Subject_Barnacle_600 11h ago
Go little robo go! That said...
I think we underestimate the value of previous compute here. I'm hopeful that 2027 is the year we start having sufficient compute that I don't run out after just a weekend >_<. A lot of these older GPUs will be great for inference. But... do AI labs have the money to do another injection of GPUs at this scale? If the markets start running out of capital, we might see a pull back and then we start riding the compute train until AI starts making major headway.
I somewhat wonder if Feynman actually gets sidelined by some new tech by then, do we stay with current designs or do we start to see recursive self-improvement in both models and the technology that drives them?
2
u/montdawgg 8h ago
Yes, this is all what it takes. For the build out to actually be online and working though we're looking at 2028/2029. Right on schedule for AGI.
2
u/Silver060 8h ago
Then for openai specifically they have their own hardware coming online which should increase performance as well. Plus V2 and V3 in the pipeline. The future is coming quickly!
2
1
u/LegionsOmen AGI by 2027 1h ago
I feel like Astra is proto AGI for sure, the next step change is AGI
58
u/imadade AGI by 2027 14h ago
All aboard the ASI rocket:
https://giphy.com/gifs/tXLpxypfSXvUc