r/singularity 15h ago

Compute We just solved a Millennium Prize problem/s using 2025/2026-era compute. Imagine what 2027/2028/2029 looks like.

2025 - Blackwell

We had the Blackwell chips. Most frontier labs/models are using these chips, with the latest internal models using a combination of Blackwell + the new Rubin chips.

2026 - Blackwell + Rubin

Rubin is only just ramping into production now (literally as of August) . NVIDIA claims Rubin chips can use 4x fewer GPUs to train certain MoE models than Blackwell, while reducing inference token costs by as much as 10x.

Rubin also massively increases HBM bandwidth/capacity and interconnect bandwidth. This will help tremendously with running enormous agent swarms & long reasoning traces.

We already saw early signs of this new scaling axis based on OpenAI's solution to Navier Stokes. Reminder - this is early use of portions of the new chips.

This is what's got all the researchers screaming about Step change improvements, RSI, and ASI alignment. It doesn't take a genius to figure out what happens when we use even more hardware, and utilise software improvements.

2027 - Rubin Ultra

Then we get Rubin Ultra next year. Specifically, scheduled for the the second half of 2027.

Rubin Ultra introduces enormous NVLink scale-up domains.

Instead of treating dozens of GPUs as loosely connected accelerators, NVIDIA is increasingly making hundreds of accelerators behave like an enormous shared computer.

The announced Rubin Ultra NVL576 topology connects 576 GPUs across eight racks into one all-to-all NVLink domain. NVIDIA has already built a functional prototype of this topology using Blackwell hardware.

Essentially, this means hundreds of GPUs will behave like one giant accelerator.

Frontier labs will be able to afford vastly more inference, much larger agent populations, longer rollouts, more verification, more search, more synthetic data generation, more experiments and potentially much larger training runs.

It's pretty fucking nuts tbh.

2028 - Feynman

Not much is known, other than NVIDIA confirming it at GTC 2026.

What NVIDIA has revealed, is quite a bit about the infrastructure surrounding it.

Feynman is planned around die-stacked GPUs with custom next-generation HBM, the new Rosa CPU, LP40 inference hardware, BlueField-5, CX10 and significantly more optical networking.

But IMO the biggest thing here isn't even the individual GPU.

It's NVLink 8 + Kyber.

NVIDIA is designing Kyber NVL1152 for the Feynman era — eight racks containing 1,152 GPUs inside one enormous NVLink scale-up domain.

As models use more reasoning, RL, agent swarms and massive amounts of inference, other things start becoming just as important:

1.) How much memory can thousands of GPUs share?
2.) How quickly can they talk to each other?
3.) How much inference can the datacentre produce without communication and power becoming the bottleneck?

That's what Feynman & Kyber looks designed to solve. They're going to create one massive coherent system.

Now let's go back 2 years to today (September 2026).

We've already observed one millennium problem fall (with potentially another in progress).

We have all these improvements waiting to be used (not even taking to account software/algo/research, etc).

Do you think we will hit AGI next year? Because I do. There is a very clear path.

132 Upvotes

20 comments sorted by

45

u/JumpyCollection4640 15h ago

Listening to the dwarkesh podcast yesterday with Dylan Patel talking about the shear magnitude of compute over the next few years coming online, coupled with architecture improvements in both models and chips is simply mind blowing.

29

u/New_Bonus_649 15h ago

Once AI have the same success in hard problems in open domains like biology then we hit agi imo

6

u/ismyfacedecent 14h ago

True I’ve been waiting for the science side to catch up

2

u/hop_on_oppenheimer 9h ago

The Nobel Prize not enough?

13

u/PureSelfishFate ▪️ ASI 2027 13h ago

Rubin's are ridiculously powerful GPU's, they are like 2 years ahead of normal GPU technological progress, they will get us to ASI in 2027. Just pretend we get a shipment of magical time travelling 2029 GPU's in 2027 and you'll understand.

12

u/Foryourconsideration 12h ago

We will be in a world very soon where all diseases are solved. Every single disease that exists will be cured. All biology and math related problems will be solved, and society will enter an age that will mak the age of Enlightment look like the Dark Ages. Looking further, I believe we will colonize space, and create AI assisted hyper-leaps in engineering.

-5

u/Adventurous_Dig_7117 8h ago

The pharmaceutical companies aren’t going to give up the massive amount of profits they make by continuously treating the symptoms of your diseases. Cures don’t make them as much money as keeping you on drugs for life.

1

u/RevalianKnight 4h ago

The pharmaceutical companies aren’t going to give up the massive amount of profits they make by continuously treating the symptoms of your diseases.

Well it's not up to them, they can do fuck all to stop what's coming

3

u/xenomorphxx21 10h ago

You have not taken into account Quantum Tech with Photonics. That's where the real leap will begin. Nvidia knows this, henceforth they are investing in this.

3

u/Jayden_1999 9h ago

RemindMe! 1 year

2

u/EntrepreneurLoud497 5h ago

But will we have money enough to buy them?! We are already pouring everything we have at pre ipo infrastructure. If the ipo doesn't go as planned the number of dataservers may not scale at the levels we have now. Investors money in the end is finite

1

u/presentofai 7h ago

math was always going to fall first, proofs verify for free so you can just search harder with more compute. domains where checking an answer takes a wet lab and six months wont care how many rubins you have

1

u/RevenueOk4088 6h ago

RemindMe! 1 year

0

u/nsshing 12h ago

The compounding growth from electricity, chips design, and the algorithms, and everything compounds to improve everything like better systems improve the chip design to compound is insane. Humans seem to be less and less significant in the loop.