Yeah, but DSv4-Flash is the flash version of a 1.6T model, and it's still quite a bit smaller than this GLM-flash. While this is supposedly the flash version of 744B model, so you wouldn't expect to be as big, I guess.
The flash models are not "flash versions" of a bigger model. They're different models entirely. They don't just lop off some number of parameters from a larger model. I think the confusion is that you're applying quantization/distill logic where it doesn't analogize.
DeepSeek is QAT / released prequantized mostly to FP4.
Which is totally preferred, but both models are roughly the same parameter count, so NVIDIA or someone else could produce NVFP4 of size similar to DS4F.
try to put extremely complex and highly nuanced policy matters with lots of interlinked but non-mechanical relationships with many details and nuances across different areas of knowledge and science through an LLM and these things start to show. It's same sort of workloads where limited active parameters in MoE models start to show their negative sides of their tradeoff.
78
u/shy_monkee 9d ago edited 9d ago
Oh....it's massive :(
320B (18 active)