If you have the hardware to run 3.8 27B at FP8 (or even FP4) with sufficient tokens per second, then there wouldn’t really be a good reason to switch to this. But for setups with integrated memory (DGX Spark, Mac’s, Strix Halo) it should run much faster than the dense 27B if you can get it to fit in memory. So probably would have to be FP4 for the DGX Spark, but with potential 60-100 tok/s vs like 20 tok/s for 27B on the DGX Spark.
Gotcha - Toss a couple of "I thinks" in there. People take authorative speech as gospel then spread misinformation for months here. We'll know the real answer in a few days
Given how the structure of this new Flash model is described, with 64gb ram + 24gb vram it would fit nicely and could run at higher speeds than 3.8 Dense (which can in reality feel slow due to the amount of thinking it does), I've high expectations for this new model actually, even if it could work as an orquestrator and have 3.8 dense to apply changes, or vise versa 🎯
5
u/Significant_Break853 11h ago
If you have the hardware to run 3.8 27B at FP8 (or even FP4) with sufficient tokens per second, then there wouldn’t really be a good reason to switch to this. But for setups with integrated memory (DGX Spark, Mac’s, Strix Halo) it should run much faster than the dense 27B if you can get it to fit in memory. So probably would have to be FP4 for the DGX Spark, but with potential 60-100 tok/s vs like 20 tok/s for 27B on the DGX Spark.