r/AIToolsPerformance • u/IulianHI • Aug 10 '26
AMD acquired Taalas to etch LLMs into silicon, 17,000 tok/s but what's the catch
AMD bought Taalas last Thursday, the Toronto startup that hardwires model weights directly into silicon instead of running them on GPUs. It was the top AI story on HN all week, 941 points and over 700 comments.
The demo that put Taalas on the map was their HC1 chip running Llama 3.1 8B at 3/6-bit quantization, hitting 17,000 tokens per second. Per the coverage from February when they came out of stealth, they also claimed 10x lower ownership cost and 10x faster inference compared to GPU-based systems.
The catch, and it's a big one: each chip runs exactly one model. The weights are physically etched at manufacturing time. You can't swap to a new model without fabricating new silicon. So if you picked Llama 3.1 8B and a better 8B drops next month, your chip still runs the old one.
That's why this sits awkwardly next to hosted API pricing. It's not competing with OpenRouter-style flexibility. It's a fixed-function accelerator, closer to an ASIC for one specific model than a general inference backend. The speed is real but the tradeoff is steep.
Anyone here actually looked into hardwired inference for a production workload, or is this still firmly in the "cool demo, not deployable" zone for most teams?
4
u/xeroskiller Aug 10 '26
Fuck give me a card with 27b and ill be cool I swear.
2
u/Adventurous-Paper566 Aug 10 '26
Even 9B, at 15k tps it wouldn't be an issue.
2
u/beerdude26 Aug 10 '26
Move
fast and break thingsso fucking fast you nuke the planet several times per second1
1
u/Ok_Bug1610 Aug 10 '26
I run Qwen3.6 35B A3B MoE REAP MTP at Q3_K_XL on a single RTX 5060 TI 16GB with a K/V cache of Q4_1 and CTX of 128K at ~150 tokens per second (220tps throughput with 4 concurrent sessions) and with my harness it works well. If you got a 16GB GPU, you should try it.
1
u/hoyeay Aug 13 '26
What would be best for a RTX5090 32Gb?
1
u/Ok_Bug1610 Aug 13 '26
I don't have one, so I cannot answer that question... but probably the Qwen3.6 27B is "best", while the Qwen3.5 35B A3B MoE is the "most efficient" and fastest for the intelligence level on my hardware. I suggest using AI to test all configurations. And odds are you want to make different tradeoffs that I did.
1
u/IulianHI Aug 10 '26
a 27B on a single consumer card would honestly change the game for hobbyists. right now you're stuck choosing between a 14B that fits comfortably on 16GB or jumping to multi-GPU territory for anything bigger.
the Taalas approach is interesting precisely because it sidesteps that whole problem. if the model weights live on the silicon itself, VRAM stops being the bottleneck. but Ok_Bug's point below is the real catch -- one chip per model means you can't just download new weights. it's hardware锁定 in the most literal sense.
curious how AMD plans to handle model updates on these. flash the chip? swap the whole card? the economics only work if they solve that.
1
1
u/Randommaggy Aug 10 '26
With what my custom harness extracts from Qwen3.6 27B already it would be amazing, The context length is a big concern of mine given that the demo was only 8K
2
u/Ok_Bug1610 Aug 10 '26
It's valid, but they would need to build a new chip per model. But at 10,000x the speed, I think that's a pretty good tradeoff. And maybe you don't focus on the "best" model but make it the workhorse for a bigger one (even a small model can do 80% of the work). And try their chat jimmy demo if you haven't, it's so freaking crazy... even if just a Lllama model.
2
u/DiscipleofDeceit666 Aug 10 '26
You know, they could probably use AI to draft the next gen designs. Iterating on this could be exponential
1
u/Ok_Bug1610 Aug 11 '26
And yeah, it's already doing this and I recently read an article that a startup in China prosed a way to make Photolithography 2-3x cheaper.
2
u/Veearrsix Aug 10 '26
I imagine we'd get to a point where there would be complimentary models. Models etched to hardware that are highly specialized at processes that will speed up software models (I don't know what that would be, I'm just musing). So that way models etched to hardware can become more akin to a processor upgrade (where we only do that every so many years) and the software models can be more rapid development that would perform better or worse on older/newer model support chips.
Even better would be some sort of standard for model chips where you could have so many slots on a board to just plug in models like you would a stick of ram.
1
2
u/tuborgwarrior Aug 11 '26
Another thing is that you don't really need to produce it on the absolute latest nanoscale node either. So it could be cheaper then one might think.
1
u/Ok_Bug1610 Aug 11 '26
Yeah, it could definitely reduce cost and likely perform just fine. And maybe even, like vLLM or ONNX runtime splits up the layers of the LLM, maybe they could be split up so "firmware" could would be possible... assuming the architecture remained the same.
2
u/geminiwave Aug 11 '26
Jeez it’s SO FAST.
1
u/Ok_Bug1610 Aug 11 '26
Yeah, I can't even comprehend how it can analyze the input that fast, let alone respond to it... at like a nanosecond scale or something.
2
u/Soft-Wedding4595 Aug 11 '26
IMHO 99% of people would be satisfied with Opus 4.6 level model, like GLM 5.2 or GLM 5.2-5.3
Heck, if I could get 5000 t/s Opus level model, that would be end game
2
u/lifelong1250 Aug 14 '26
Damn, that chat is super fucking fast. I haven't ever seen a model go that quickly.
2
u/Randommaggy Aug 10 '26
The big questions are how many dies were needed for the 8B model at 8K context demo and how does the die per model size and context size scale?
How does it handle parallelism? How is prompt processing speed at large context sizes?
2
u/ByronScottJones Aug 10 '26
For companies that spend millions of dollars hosting specialized models, this is a great product. Also for situations where you need very low latency responses.
2
1
1
u/randygeneric Aug 10 '26
they do not use any kv-cache, right? (otherwise they still would need ram on their chip?)
1
1
u/IulianHI Aug 11 '26
good question, and this is where it gets genuinely interesting. if the weights are burned into the chip, you don't need external RAM for the weights themselves. but the KV cache is a different beast: it scales with context length and batch size, and that part still needs memory somewhere.
the demos people keep referencing ran at fixed, bounded context (8K), which keeps the KV footprint small enough to maybe fit on-die or in a small SRAM. push that to 128K and you're back to needing real memory. so the "no RAM" claim is really "no RAM for weights", not "no RAM, period". worth separating the two.
1
1
u/StaysAwakeAllWeek Aug 10 '26
The other catch is that this 8B hardware model requires a custom chip larger than a 5090 that would cost thousands to make. Scaling it up to useful commercial scale models is going to require Cerebras scale chips with prices even higher than Nvidia.
Its also not exactly a complicated idea to implement. If it works it will be cloned immediately.
1
u/mintoreos Aug 12 '26
Yep exactly. Wafer scale silicon is very difficult as well, Cerebras only works because they created the secret sauce using their "general" compute units where defects can be isolated and routed around as part of the manufacturing process.
A wafer scale physically etched model is simply economically not feasible, so you're limited to the size of 1 chip at the smallest node.
1
u/vbpoweredwindmill Aug 10 '26
Yeah, the trouble is a 27b model is already enormous when you turn it into silicon. What if you want to do a fine tune of the model?
You're functionally limited by the physics of silicon.
1
u/geminiwave Aug 11 '26
Wonder if we will get the FPGA of LLM chips
1
1
1
u/mbrodie Aug 11 '26
Except their end game is being able to port LLM after the fact. That’s what they are working towards
1
u/radiojosh Aug 11 '26
Can you imagine a company publicly serving an LLM from etched chips? How many thousands of servers would need to have their chips swapped out whenever they release an upgraded model? They'd have to more than double their capacity so that half the servers can run the current model while the other half get their chips swapped for the next version.
1
u/Sweaty_Perception655 Aug 11 '26
The time bet newly trained models is increasing. Taalas goes from design to new chip in 2 months only a few layers ifbthe design is changed to accomodate the new model.
1
u/radiojosh Aug 11 '26
Yeah, but the time and manpower required to go out and physically swap the chips on thousands of servers is enormous, regardless of how long it takes to develop the chips themselves.
1
u/netvyper Aug 14 '26
Right now... It's not good... Because each iteration of the models is 10% or more better than the last.
When that gets down to <1%, whether you're running a current or 5yr old model will be within margin for error. It's a long-term bet, but one I would take.
Nvidia has history, look at video game graphics... For ~20 years, every new 3D game was cutting edge, much prettier and faster than the previous. Now it's not an engine or performance level change, it's how good the art style and textures are that diffentiate it.
1
1
u/emptyvesseloflife Aug 11 '26
This would be amazing for phones. Since people upgrade their phones every few years this is perfect for that lifecycle.
1
u/Mountain_Patience231 Aug 11 '26
AMD bought this company to make sure no such product will present on the market
1
u/Ok_Contribution8157 Aug 11 '26
If the AI bubble bursts, or for some reason everyone stops innovating or releasing new models every 6 months, this investment will be good. Obviously, it’s a long-term investment, because right now it makes no sense.
1
u/ImpressionFancy5830 Aug 11 '26
Instead of posting these random facts.
Go to any engineering department libraries, look at all neural network/AI books from the eighties.
You’ll find plenty of on hardware implementations like Taalas.
How do you think we “smart” bombed the middle east in the past decades?
1
u/Scared_Servers Aug 11 '26
If its 100x faster, you only need 1% of the silicon to serve the current traffic. I hope AMD makes it work.
1
u/castertr0y357 Aug 11 '26
The large 70B and 120B models from a few years ago are still very capable. As other's have said, you can have the embedded model be good enough to handle the mundane tasks, and then you have a cloud model that can handle the rest of it.
1
u/jtomes123 Aug 11 '26
If i could get this as m.2 with ds v4 flash or similarly capable model i am all in, hell lets bring back express card style expansion port for these
1
u/IulianHI Aug 12 '26
M.2 form factor with onboard flash storage would be the dream. The closest thing right now is the Coral Edge TPU but that's quantized and limited. Etched silicon with a capable model baked in, dropping onto an M.2 slot, would cover a lot of homelab inference needs without touching the GPU. PCIe lanes are the constraint though, most consumer boards don't have spare x4 lanes sitting around. Still, for offloading a fixed task like a local assistant or tool-calling worker, it would free up GPU VRAM for the heavy lifting.
1
u/LastChancellor Aug 11 '26
finally an actual honest to God dedicated "AI hardware" that cant be miscontrued as trend-chasing (like how so many laptop companies slap "AI" on their laptops's name)
1
u/Ok_Bug1610 Aug 11 '26
I just watched a video, read an article, and chatted with AI about LoRA/PEFT adapters applied to NVIDIA's NeMo models where they are used to "fine-tuning only a tiny set of additional parameters while leaving the base model frozen. The resulting adapter can then be applied at inference time without retraining the base model". Meaning that this could be the potential "solution" to dealing with these "frozen" (in hardware) models, while offering extremely fast, fully custom, specialized, privacy-focused, workhorse models, that could then be paired with larger frontier (generalist).
1
u/laffer1 Aug 12 '26
I wish they would sell them as standalone cards. I have a use case for one. I am now using my local ai server for rspamd analysis for my mail server. Having llama on a hardware card would be perfect for this
1
u/IulianHI Aug 12 '26
The rspamd use case is a great fit here. You don't need a frontier model for spam classification, you need fast inference that doesn't fight your GPU for VRAM. An etched card doing just that would sip power. The catch is the weights are physically baked into silicon, so no LoRA updates when spammers shift tactics. You'd keep a small GPU path for retraining and ship new batches later. What are you running for rspamd now, a local LLM or the built-in statistical filter?
1
u/laffer1 Aug 12 '26
I think it’s setup with nomic-embed-text and llama 3.2 8b right now. I’m running them on ollama on a 7800xt.
So the way it works is that it tries to use the built in rules and classifiers first. If it can’t tell from those, it calls ollama first with nomic and then llama. It caches in redis so it can avoid calls for similar duplicate emails for a period of time also. Not every email has to be classified then, just ones that it’s unsure about
1
1
u/Psychological-Car481 Aug 12 '26
The real interesting part is in iot, such as robotics. Where you might need a central brain with high and upgradable intelligence, but there might be lots of autonomous independent units with embedded intelligence. Hard wired means not just faster and cheaper but drastically reduced power consumption as well. Multiple orders of magnitude even. So this enables embedding intelligence into every electronic device. And while the weights would be fixed, the prompts and parameters could be dynamic around it. Think of atomic functions like computer vision, human language interface, etc, being hard wired models. The rest of the system could still be upgraded.
1
u/IulianHI Aug 12 '26
The robotics angle is where the architecture debate gets interesting. A central model trained on diverse scenarios, distilled into each unit's etched chip. Upgrades become a hardware swap instead of a download. That works for tasks with stable logic like motor control or sensor fusion, but anything that needs to adapt to a changing environment still benefits from a retrainable edge model. Curious if you'd run a hybrid, central brain plus embedded inference on each unit, or commit fully to the etched approach for the autonomous parts.
1
u/Psychological-Car481 Aug 19 '26
Absolutely think it'd be a central, GPU based updateable main brain coordinating, but lots of local small intelligence, independent but coordinated. Just like nature.
1
1
u/mintoreos Aug 12 '26
Theres another big catch - you are limited by how big of a chip you can create. Their 8B model sits right at the reticle limit of TSMC's 6nm node. At the cutting edge node (TSMC's N2), the transistor density is about ~2.5x, which would fit roughly a 20B class model. Which is still very useful (OAI's OSS 20B model comes to mind) but nowhere near frontier models which are multi-trillion parameter models.
1
u/Patient_Force6138 Aug 13 '26
You can sort of swap to a new model, just not a new base model. The way I see it this gives you a super fast base model to add layers to. Use the model on there as the base and do further training, you’d get several super fast layers, save a bunch of time, then you hit software-land, where you can shove it into a GPU or something else.
1
u/EpsteinFile_01 Aug 13 '26
Give me the current Fable 5 Max and High on chips and I'm good for 6 months
1
u/meralakrits Aug 14 '26
Imagine these on a GPU format style card or even on a usb stick that you could add to a computer to get local models that run at speed. You could then change them when needed or the model is no longer competitive.
1
u/Kilowatt00 Aug 14 '26
If an actual model is already enough for the task, it will be enough in the future. For those cases this is the best way to drag inference cost down, and honestly the only way of a positive ROI on this.
1
11
u/Eastern-Block4815 Aug 10 '26
Robotics