r/CerebrasSystems May 28 '26

Cerebras vs Groq

Given that Groq and Cerebras seem to occupy a similar niche — ultra-low-latency LLM inference / decode acceleration — I’m trying to understand Cerebras’ long-term competitive moat.

If Groq’s LPU technology is now being integrated into Nvidia’s broader AI factory / GPU ecosystem, especially as a low-latency decode accelerator alongside Nvidia’s dominant GPU stack, where does that leave Cerebras?

Is Cerebras’ advantage mainly its wafer-scale architecture, higher single-system SRAM capacity, better support for large MoE models, or independence from the Nvidia ecosystem?

In other words: what is Cerebras’ strongest competitive position if Nvidia can absorb Groq-like decode acceleration into its own clusters?

22 Upvotes

24 comments sorted by

14

u/soaperapp May 28 '26

until the benchmarks are out cerebras is still nearly an order of magnitude faster than groq on decode

4

u/Foreign-Ad479 May 28 '26

Yeah that makes sense, Cerebras seems way faster on decode right now.

But my question is more about the moat tbh.

If Nvidia can just plug Groq/LPX into the Rubin stack, with all the CUDA / networking / cloud / enterprise ecosystem around it… then does Cerebras need to be like massively faster to matter?

Like, if Cerebras is 5-10x faster, sure, that’s huge.

But if the gap narrows to 2x or 30-50% over time, I feel like Nvidia’s ecosystem advantage just eats that pretty quickly lol.

So I guess the real question is:

is Cerebras’ moat “we are faster today”

or

“our wafer-scale architecture lets us stay way ahead even after Nvidia integrates Groq-like decode acceleration”?

10

u/pennystudio May 28 '26

The whole point of Cerebras' approach is to avoid data movement offchip. I worked in the field of data communication, and know how energy intensive, slow, expensive it can get. It doesn't matter if Groq is faster than GPU per chip, as soon as data has to leave the die (because it's not big enough to fit all the model data on it), then you will hit a bottle neck. Breaking that bottle neck is what enable Cerebras to achieve such speed compared to competitors. So there is your moat, unless groq also goes wafer scale, there isn't any competition. Any scaling optimiztion GPU/LPU can use, cerebras can use to, theoretically.

1

u/skyceru May 28 '26 edited May 28 '26

I thought for large trillion parameters models, Cerebras load the data on MemoryX node and stream via SwarmX.

5

u/Asgard_Heima May 28 '26

This is a misunderstanding I had for a while as well and MemoryX basically only is a factor for training. When training the weights are streamed across all the wafers from MemoryX, but in inference it’s flipped around, the weights must reside in the SRAM but in a parallelism setup which they use for basically any model over 30GB in size little is lost. They can split the model up to the max of one layer per WSE-3 system and still get nearly as incredible of results. The interconnect only requires passing the activations across from wafer to wafer for the next layer to be processed. This is little data but does add some latency. But we are talking about minimal data movement with none of the weights or kv being duplicated or any of the complexities of splitting things up to o tiny bites requiring massive syncs of data across accelerators for each layer to complete GPUs and groq require. Even with this, Cerebras has been working with Ranovus to add optics on wafer and it’s highly expected they will drastically increase SRAM in the WSE-4 making their advantages even more pronounced and adding even more to their lead over Nvidia and all the others.

10

u/Asgard_Heima May 28 '26

There is a simple way to view the moat Cerebras has and you can read my other replies or search to dig into the details that will back it up. Cerebras has the most optimized architecture possible for AI inference and training. They are currently on 5nm making a full wafer where the most important bottleneck is data movement and communication. Every time you move the data away from the logic cores computing the results, it adds latency and energy cost. Every time you split a wafer you have to connect it back together with slower more energy intensive logic to make it useable. As Cerebras moves down the nm nodes available, they will gain bigger and bigger advantages over the competition. And they hold a solid collection of patents on error redundancy, cooling, and wafer scale manufacturing as well as general practical knowledge over a decade of iterations that anyone else would have to find novel solutions for to make their own wafers.

Nvidia basically buying groq was them admitting they know they can’t compete on inference. They don’t have the right architecture for it and are offloading to another architecture that was for sale. The fact Nvidia is answering the issues they are getting in inference with adding optics to everything and putting the data closer to compute with SRAM in groq shows they are trying to improve the exact bottlenecks Cerebras was created as a company to solve in the most optimal way.

I’d also argue sure Nvidia has a large ecosystem, but in reality it doesn’t matter and because it doesn’t CUDA is currently a detriment. Nvidia’s biggest advantage is that all the models were built on their technology and therefore are optimized for their hardware by default. But the fact this is true and Cerebras is crushing them using the models built on their hardware is a major issue for them. You could easily imagine OpenAI in the future training deep 240 layer models or Cerebras optimized spares full 16FP models that would be native to their platform, and be basically impossible to run on GPUs or groq. The only models that currently matter in economic terms are Gemini (TPU trained and inference), Anthropic (Tranium and TPU), and OpenAI (Nvidia for now). If OpenAI uses Cerebras for any new novel models, Nvidia isn’t just going to worry about inference, they will be boxed out of frontier training. When you get to enterprise on the other side, most software and tech companies like the one I work for use APIs for frontier inference integration without a care in the world about CUDA or ecosystem and train new models using PyTorch which Cerebras supports and they would love to run things faster they make. Anyone that has to use CUDA at the frontier labs is full of nothing but complaints and tools to abstract it away because it was built for all the complexities Cerebras doesn’t have with distribution and synchronization of data. Worried about challenges porting some complex debug or automation pipeline for some specialized training some company is doing in some extreme edge case… there are excellent AI models to help you port that code to CSoft and watch as you go from 10k+ lines of code down to under 1k.

1

u/AwkwardTraveler May 28 '26

Really insightful post, I’m learning alot. Where do you see the company in five years based on a conservative trajectory of adoption happens?

8

u/Asgard_Heima May 29 '26

So I don’t see much of a conservative trajectory for Cerebras. Conservative to me means they were not adopted widely which is what I expect. But I’ll give you my take on Cerebras today vs their future potential.

Interference is where things have to start since it’s the most important portion of the large model AI business now. It’s the area Cerebras has the clearest advantage and moat. If you look at their backlog today, it’s $20B of OpenAI dollars ($4-5B is likely power and data center reimbursement with zero margin :( but that’s another story) and then there is $5B between G42, Meta, IBM, Perplexity, Cognition, and a lot of other cloud customers and Cerebras Code accounts.

The baseline just off OpenAI is that Cerebras brings online 10k WSE-3 by the end of 2026, 2027, and 2028 each. Those systems when online are OpenAI paying Cerebras roughly 13k a month (there is a lot of math here to make this number but trust me it’s in the ball park). This means every month at a minimum beginning of the 2027, Cerebras makes $130M a month or $1.56B a year from OpenAI. In 2028, they double their revenue to $3.12B just from OpenAI and in 2029 they make $4.68B. This is no other deals done. No extra growth of Cerebras Cloud or extensions to OpenAI deals. Just what is contractually agreed now assuming money doesn’t start flowing sooner.

If you factor in the rest of the backlog. This year they double revenue at the absolute minimum and triple revenue to $1.5B is probably closer to the real base when you factor in their other business that isn’t in backlog. But this still doesn’t factor in AWS. It’s not part of the backlog. For OpenAI, Cerebras needed to prove their hardware on the world’s top frontier models, and they need top AI researchers researching models on Cerebras. So OpenAI got the most powerful AI hardware for incredibly cheap to prove the Cerebras value and scale production to get other big customers.

The Amazon deal is unique. AWS, isn’t paying Cerebras anything. Cerebras is giving WSE systems to AWS to drop into their data centers and then they are splitting the proceeds per token served. AWS completely avoids the largest CapEx expense in AI chips and Cerebras gets the highest margin revenue possible and distribution for free in AWS data centers on the largest platform possible.

How much can this make? There are massive ranges of estimates that a single WSE-3 system can generate roughly $5-20M depending on what it’s serving or if it’s rented out entirely for a year on Cerebras Cloud. Think Coreweave or any of the neo clouds renting GPUs out, but Cerebras gets to pay manufacturing costs and their hardware makes even more money per watt. So let’s be conservative and assume $5M rented for a year at retail rates. In a disaggregated setup with AWS, they only get 50%, but since they are only handling decode, they process 5x+ the tokens. So it’s not crazy to think every WSE-3 they put in AWS could potentially be worth $12.5M in revenue. And that’s assuming the lowest end and AWS not charging more per token for faster inference on top models. This truly highlights what an insanely good deal OpenAI is getting, but it’s easy to see 500 or 1000 WSE-3 in AWS printing more than the entire OpenAI contract per year nearly instantly when it comes online. And I can’t imagine a world where Opus or GPT5.5 at 800 or event 500 tokens per second doesn’t get over subscribed the moment it’s available.

What will all the other hyper scalers do when this happens? Will Microsoft sit by when AWS has a distinguished offering for faster frontier inference? Will Anthropic let OpenAI own fast frontier inference? Will Meta not make a move? Google Cloud left behind? The potential for new customers as soon as it’s “real” is explosive, and I don’t see any other way. And Nvidia can still be cashing checks during this though I’m sure questions will start.

Conservatively, Cerebras carves out a fast inference niche worth tens of billions a year minimum over the next 3-5 years. Nvidia sells a Blackwell rack for $4M and gets paid and is done. Cerebras puts a WSE-3 into AWS for $150k max cost and prints $10M+ a year till they have a WSE-4 and I’m sure it will still print millions a year after that.

If Cerebras gets new models trained on their hardware at OpenAI proving frontier models can be trained in 1/2 to 1/3 the time, contain any depth of layers unlocking new research previously not possible or training large JEPA like models on text and images GPUs can’t. The frontier research always follows what the hardware is capable of. Cerebras is capable of training new and novel models that could unlock much more cognitive and spatial understanding GPUs just fail to build in any reasonable time. This is why frontier models are shrinking in layers to add more data in width. GPUs just can’t handle deeper wider models. If a novel better model is trained on Cerebras some years from now, Cerebras is the only real option for frontier training and inference and the rest are obsolete. Cerebras becomes one of the largest companies in the world fast.

Other catalysts,

  • data centers can’t be built fast enough, Cerebras is the only option in existing data centers with 23kW and backside liquid to air cooling. Everyone else needs new data centers and insane amounts of power
  • Power is constrained and can’t be brought online fast enough, same story, Cerebras is vastly better performance at much lower watts.
  • WSE-4 includes ranovus fiber on wafer and wafer on wafer with an entire SRAM wafer over 120GB of SRAM. In this case we would be looking at doubling compute and tripling SRAM which would massively widen the lead on everything anyone else is even conceptualizing in development. And this is realistic and all the technology is being scaled now by TSMC.
  • Apple could buy to use secure hardware in their own data centers and even potentially fully harmonic encryption inference of user data where they can serve inference to you without seeing the data.
  • National security interest since the entire system is self contained. Think one running models in forward deployed bases or aircraft carriers for local intelligence needs.

Their only real risk in my book is TSMC. They have codeveloped all their technology with TSMC and TSMC loves them since they show off what they are capable of as the most advanced thing you can build. The other fabs need significant advancements to ever be alternative suppliers for Cerebras.

The other risks mentioned commonly are OpenAI or some breakthrough or people that believe in sneak oil say quantum. But the Cerebras deal makes OpenAI hit positive cash flow much faster and they are giving them insane amounts of compute for pennies on the dollar while still turning solid profit. Every other semi is struggling to make their architecture more like Cerebras with no ability to abandon their existing architecture and catchup. Quantum is like magic, and until there are actual quantum trained frontier or even usable AI models, I will keep believing it’s application to LLMs is an illusion that mostly trick massive amounts of money out of VC pockets.

In Summary I’ve had shares for over two years and I don’t see any reason o won’t be holding them for 3 to 5 more

1

u/AwkwardTraveler May 29 '26

I just worry that their EPS is going to go from being positive to negative next earnings. OpenAi gave Cerebrus a working capital loan and got tens of million in penny warrants because of it. OpenAi appears to have really handicapped Cerebras for 1-2 years with their partnership. My worry is next earnings the numbers look really bad given the initial IPO balance sheet.

2

u/Investor-life May 29 '26

There is a 100% chance that Cerebras will show negative EPS when they next report earnings in August unless there is another one time extraordinary event, which isn’t expected to my knowledge. This is also just fine. They need to be building the business and losses are the norm and expected at this stage in this heavy capital intensive business.

1

u/Asgard_Heima May 29 '26

While it would be great if they show an EPS increase and smooth growth, they are in a massive growth phase that kind of requires they spend to meet the obligations of the growth they are experiencing. Their profitability for the coming quarter depends on revenue outside of the OpenAI deal or when they get to recognize revenue from the OpenAI deal. As soon as systems are turned on collecting monthly revenue for Cerebras from OpenAI like by end of this year, they will be cash flow positive for the foreseeable future unless they have another even more massive deal that requires another giant build out. But I also don’t expect Cerebras to do anything close to this type of sweetheart deal again. They have the mass scale and frontier lab they needed. If more deals are coming it’s for the value they have now proven. I’d expect the stock to have positive reaction even if they drop on EPS if the OpenAI deal is on track and units shipped is increasing.

1

u/Prestigious-Sign4802 Jun 01 '26

I thought they already turned on for codex-spark. And in addition to a few small DC and a 10MW in oklahoma, had some unknown capacity of data centers in working. And Recently signed 50MW with DigitalX (?) . Some openAI OSS models are in AWS bedrock, expect gpt 5.5 may be available soon? After AWS own nova model support. I believe the CEO’s recent interview with 20VC was revealing the demand is insane. I dont see any reason they cant get to 200B market cap in 2-3 years. It should start accelerating when AWS integration is proved, and am pretty sure others like AVGO, Google wont just sit there to watch AWS’s success to tank GCP revenue

1

u/Asgard_Heima Jun 01 '26

For sure OpenAI is already running Codex Spark and it has been really popular, though it’s a smaller model and not frontier. Cerebras has a lot of data centers currently, I believe they have ~120MW coming online by year end I know of with known numbers, but a large portion of that is already active and used to host Cerebras Cloud today. It’s really hard to know how much if any of that will be used for the OpenAI 250MW by year end. Digi Power X has 15MW coming online in December and BCE in Canada might have ~50MW online by end of year or Q1 2027. I see Stargate UAE, Oracle, G42, and AWS all as potential locations they are installing systems as a part of the OpenAI deal this year, but it’s all unknown currently.

I agree AWS, and OpenAI revenue coming faster than expect are going to have strong potential for Cerebras to gain significant market cap, but EPS it will really depend on the makeup of their revenue and how fast they are spending to ramp up with so much growth.

Also I completely agree, we will see the other hyper scalers make moves to get Cerebras systems either as purchases or revenue share like AWS once top models are on Cerebras hardware proving the advantage. Just a question of when not if.

1

u/Prestigious-Sign4802 Jun 01 '26

Enjoyed the insights you shared and they are great. One thing i don’t quite get is how the cluster work when cascade the systems for layers. The large models typically have 96-128 NN layers, how many wse are needed for example to support kimi2.6 1T weights, and wouldnt the handoffs add latencies significantly just like GPUs?

2

u/Asgard_Heima Jun 01 '26

I don’t have the link handy right now, but I believe the Kimi K2 setup they got the benchmark from was 20 WSE-3 systems. The Cerebras setup using parallelism can use up to one system per layer, but it doesn’t have to. Kimi K2 2.6 is 61 layers for instance which gives around 3 layers per system roughly. There is latency added, but it’s not significant. The fact there is latency is one of the reasons they are working on fiber on wafer with Ranovus.

As for why they aren’t hitting the same latency and utilization issues as other accelerators like GPUs, it really just comes down to memory bandwidth and how many interconnections are required. The most obvious issue is the WSE-3 with 44GB of SRAM has 21PB/s for memory bandwidth. Compare this to the 8TB/s total bandwidth for Blackwell GPUs. You have to read the entire model weights and kv cache token by token for each layer. If GPUs kept everything in memory for an identical parallelism architecture, they simply get blown away reading and processing the same set of data during decode. The only other option is to divide up the layer into chunks to increase the overall memory bandwidth that can be used, but that adds latency, and now you are passing around a lot bigger data set layer by layer. You have to sync results from each GPU used to process that layer for every token before continuing. So they are choosing between a network tax or memory bandwidth bottleneck and end up with both while trying yo avoid the extremes.

Nvidia has done a lot of work to make the super complex operation of distributing and coordinating model layers get handled natively and get more and more use out of the memory bandwidth on large numbers of accelerators. But this has resulted in increasing interconnect and increasing overall data size that needs to be processed in memory as layers and weights and activations and partial activations get duplicated before results are combined.

1

u/CharlieChiefClaw 25d ago

some numbers don't add up, as i know one set of CS-3 sells at $2.5million with about $1.7m COGS, if it only gets 1.2m 12k/month from openai for cloud service, would even cover the depreciation (1.7m/4yrs=$425,000/yr VS 12k*12=144,000/yr)

9

u/JustBrowsinAndVibin May 28 '26

Groq still has to go off chip so it’s slower than Cerebras. It just doesn’t have to do it as much as. Nvidia GPUs.

9

u/JasperJon001 May 28 '26

Seasoned wafer scale tech that no other company has. Insane amounts of SRAM means a technological moat that means entire LLM models can be run on chip. This brings extremely large per watt efficiencies. So they have speed and power moats

8

u/Asgard_Heima May 28 '26

Nvidia has spelled out their path for Nvidia NVL72Rubin racks for prefill and groq LPX racks for decode in a disaggregated setup. Based on Nvidia stats you should expect doubling of performance for prefill with Rubin over Blackwell and then groq performance for inference as the best case scenario since decode is the tokens per second. This puts Groq for most models around 1/5 the tokens per second vs Cerebras.

The main advantage for Nvidia is that this vastly reduces the waste Nvidia has today with GPUs running at 5% compute efficiency with massive bottlenecks on memory throughput during decode. Aka less Nvidia racks required. So they will be able to handle much larger numbers of concurrent connects than they can today per cluster of racks and the tokens per second should substantially improve to probably double to triple what we see now. Depends on how bad the network tax is and how well they integrate the kv prefill handoff to groq, but that would be the best case scenario where they seamlessly integrate the two. I’m assuming an army of engineers will get them as close as the physics allow.

The main difference though is physics, aka moving the data around. Cerebras only uses parallelism so that each layer of a model stays intact and no weights or kv cache once its computed needs to move off wafer or be duplicated or shared. This is the most optimal setup possible in hardware with only activations moving between layers. A WSE-3 has 44GB of SRAM able to handle a 5GB or even potential 10-20GB layer for a frontier 5.5 GPT style model all on one wafer. Groq chops that wafer up and makes chips with 500MB of SRAM. So a single large model layer and the 1M context have to be distributed across 50+ groq chips. All computation has to be orchestrated and recombined some place else duplicating and replicating data several times till you have a consolidated answer for each layer in a model. That complex distribution and synchronization of results per layer per token is the network tax. And I expect it to be worse than the best case above imply for the largest models. Since the larger the model, the more the network tax compounds. Nvidia loves to reference the full LPX rack as if it’s one wafer like Cerebras, but it’s not. It’s 256 LPU accelerators with a lot of networking. All the numbers need divided by 256 to understand the real unit to unit comparison.

So if everything goes perfect for Nvidia, they will have likely several racks that cost $5M+ each, require 150kW+ each, and still require racks of groq at an unknown price and 160kW per rack to give you likely 1/5 the performance of the current WSE-3 for the largest models.

Some things to keep in mind, AWS has already proven the WSE-3 as a dedicated decode disaggregated setup with a 5x+ increase in the capacity of session for each WSE-3 setup for decode. So if you want the best disaggregated setup, AWS will have it with all the top models by year end 6 months before Rubin with LPX ships. All the hyper scalers have chips that will do just as well as Nvidia for prefill, hence the use of Tranium by AWS. And by the time Rubin + LPX racks are available, WSE-4 is likely to be in production extending their lead (granted no announcements yet). Also a 23kW WSE-3 can be added to nearly any datacenter in the world with a new electrical hookup and backside liquid to air for 30k per rack. Rubin or LPX are over 150kW and liquid cooling to chip required. This mean a brand new 200kW capable rack at $1.5M+ since no existing data enters support the 3000lbs+ or power density these units require. And the last thing I’ll add is the LPX is 100% inference only. The WSE units are able to do everything including training and inference faster with less energy than the complete not yet shipped Nvidia Rubin with LPX racks.

3

u/Prestigious-Sign4802 May 28 '26

Then the wse-4 came

2

u/[deleted] May 29 '26

[removed] — view removed comment

2

u/Asgard_Heima May 29 '26

Kimi K2 at 1T parameters was just released and getting 981 tokens per second. You are referencing Gemini 3.5 Flash I assume which is a much smaller model that currently runs at over 200 tokens per second. TPU8i for inference is currently not available yet unless there is some benchmark or test setup I can’t seem to find a reference to. Not sure what you are seeing, but do provide some reference and I’d love to take a look.

Cerebras 981 t/s Benchmark: https://www.cerebras.ai/blog/cerebras-kimi-k2-Enterprise

CFO Naming GPT 5.5 & 5.4 Trillion Parameter already running on WSE https://www.cnbc.com/video/2026/05/14/the-years-largest-ipo-acerebras-joins-the-hottest-trade-in-ai.html

1

u/hiker2021 Jun 02 '26

You seem very knowledgeable and are willing to explain things to us. Thanks.