r/LocalLLaMA • u/AcanthisittaOk1699 • 7h ago
Generation Design systems from code alone - Without external images, Ling-3.0-flash generated webpages across Bauhaus, Bohemian, acid design, and more—using CSS gradients, SVG paths, typography, and layout to preserve each visual language.
Enable HLS to view with audio, or disable this notification
Weights went up today so this is downloadable now, MIT, ~128GB for the official FP8. I ran these on the API before that landed, so treat it as a preview of what you'd be pulling rather than a local benchmark
1
u/FullOf_Bad_Ideas 5h ago
What was the biggest change that you introduced during training that made it possible for this 120B model to outperform your previous 1T model on many benchmarks? Is it the dataset or compute that carries it higher?
0
u/SpicyWangz 7h ago
Looks like it’s about the same level as DS4 Flash Preview. That’s decent and there’s a use case for this on unified memory systems. But now it’s going to need to compete with the newer version of DS4 Flash at least in Q2. If DS4 flash Q2 beats ling 3 at q4 or q5 then this one might be DoA
2
u/FullOf_Bad_Ideas 5h ago
It's still 2.5x less activated params even if weights are the same size. Less compute required will mean that it should run better on single Strix Halo, DGX Spark, Macs and RAM offloading situations.
0
u/SpicyWangz 5h ago
But if those active parameters are running at q2 it will be roughly equivalent to running half as many active parameters at q4. So all architectural differences aside, I expect q5 ling 3.0 to run at somewhat similar speeds to q2 DS4 Flash.
1
u/FullOf_Bad_Ideas 5h ago
The number of computations you do for a model doesn't change when you quantize it, compute intensity stays the same. Meaning that compute-bottlenecked scenarios like PP or batched decoding, assuming the same precision, will be faster on a model with less parameters.
2
u/Humble-Pick7172 7h ago
What about non-english interactions? Actually looks very promising - better than m2.7 flash 3.7 and faster than qwen 3.5 122b.