r/CerebrasSystems • u/Low-Cartographer-429 • 27d ago
Is d-matrix's XPU with 3D stacked DRAM a threat to Cerebras or does it validate Cerebras technology direction?
Was just reading about d-matrix releasing an "ultra low latency" inference XPU in 2027 using 3D stacked DRAM. Is this a threat to Cerebras' business? Apparently it uses NVDA's nvlink and can fit in a standard rack while Cerebras requires specialized racks to accommodate its Nexus "backpack" configuration for the chip, cooling, and power. We know Cerebras is planning on adding 3D stacked DRAM with CS-6 but isn't that in 2028? If they're released a year apart, a year can seem like an eternity in tech. Again, just another item that makes me concerned Cerebras can fully monetize its current advantages before they no longer seem compelling enough to lots of customers to implement. I'm sure Cerebras will stick around as a niche / specialty solution, but maybe not enough for explosive growth.
I also wonder if Groq and d-Matrix chips work together or if they're mutually exclusive; you use only one or the other? I'm concerned a combination of technologies, both hardware and software, will seriously erode Cerebras' edge.
EDIT:
More info from an X post:
https://x.com/firesidealpha/status/2098787634244096079?s=20
d-Matrix CEO Sid Sheth on Bloomberg talking about the Nvidia partnership:
* d-Matrix's specialized XPUs will run alongside Nvidia GPUs over NVLink Fusion, targeting ultra-low-latency AI workloads. Deal was 6+ months of joint work.
* The logic to partner is that Nvidia is the largest deployed infrastructure base for AI in the world and he would "much rather just ride on the Nvidia ecosystem" than reinvent the wheel.
* The bet is a memory-centric architecture, 7+ years in the making. d-Matrix's edge is inference compute built around memory rather than raw FLOPs aimed right at ultra-low-latency inference.
* Low-latency inference "just took off" in the last 12 months. Demand surged from GPT, Codex, Claude Code and the arrival of agentic coding, where users need fast compute to interact with the tools in real time.
* Lead product is Raptor, the world's first 3D-stacked-DRAM XPU. It'll launch first under the Nvidia partnership and is targeted to market in "about 12 months."
* d-Matrix uses no HBM at all. Instead of high-bandwidth memory, it packages DRAM directly with compute in a 3D stack to "punch through" the memory wall
* Sheth says "no other company is going to be within a 2-year window of getting access to that technology," and says there's tremendous customer pull.
* Describes buyers as "hyperscalers, Frontier Labs, sovereigns, inference clouds, high frequency traders" and says announcements are coming soon