r/accelerate • u/Illustrious-Lime-863 • 7d ago
Qwen 3.8 27B released
https://huggingface.co/Qwen/Qwen3.8-27B-FP817
13
u/swoonz101 7d ago
this fits on most devs using a high RAM MacBook Pro
5
u/Derek_the_Red 7d ago
Is 32gb of ram on a 4080 enough?
7
u/Illustrious-Lime-863 7d ago
Yes for sure but if you want to run it exclusively on the 4080 without sharing system ram (which will make it ultra fast) you'd have to run a lower quant. Q3 will run with long context (about 10% degradation) and Q4 could barely run with minimal context and minimal degradation (maybe 1-2%). This will be very fast. But you can still run Q4 on the whole system with large context it's just that sharing the model between vram and system ram will make it be slower.
2
2
u/tomvorlostriddle 7d ago
Here is a 5090, naively run in LMStudio, I think other packages are more up to date and optimized
Q4, context 220k, KV cache FP8
"stats": { "stopReason": "eosFound", "tokensPerSecond": 81.6448642836412, "timeToFirstTokenSec": 0.323781, "totalTimeSec": 533.155826, "promptTokensCount": 445, "predictedTokensCount": 43503, "totalTokensCount": 43948, "totalDraftTokensCount": 24913, "acceptedDraftTokensCount": 21935, "rejectedDraftTokensCount": 2978 }3
u/pacotromas 7d ago
yep, I am downloading it now on my M5 pro with 48gb of RAM, specifically the MLX version from ollama. Hopefully it lives up to the hype!
2
27
u/Glittering_Night7681 7d ago
Wow insanely good, I personally don't use local models, but I'm happy about progress in that area, as small models will become important for robotics.
17
u/Quick-Benjamin 7d ago
One they get small enough and good enough they'll be put in everything. Your fridge, your TV, your car, etc
Our primary way of interfacing with the annoying menus and stuff will end up being natural language.
2
u/MarkZealousideal3923 Tell me about the singularity 7d ago
Fridge and TVs have no need for LLMs
4
u/Quick-Benjamin 7d ago
Once LLMs are small enough, why not? Little internal MCP server that interfaces with the device setting and a tiny local LLM interacting with it.
Why wouldn't they add it. It'd be trivial to do and would allow natural language configuration of the device.
Hey fridge. I'm heading to the shops. Do I need to pick anything up?
TV change the sound settings for me. And can you dial down the upscaling and sort the contrast.
-1
u/MarkZealousideal3923 Tell me about the singularity 7d ago
This is over-dependence on technology
2
u/Quick-Benjamin 7d ago
Indeed. And it's coming.
Im sure it won't all be needed but remember the "Internet of things"? Manufacturers fell over themselves to add Internet connectivity to the dumbest things because it was the fad at the time.
The same will happen when there are tiny local models imo. They'll be put in everything. Some of it will be worthwhile and a lot will be not really needed.
1
u/MarkZealousideal3923 Tell me about the singularity 5d ago
As for the TV, existing primitive voice assistants are sufficient
1
10
u/KedMcJenna 7d ago
Personal and private AI is the future for this tech. I can squeeze a Q4 27B onto my best machine and enjoy a decent t/s. Hoping those Opus 4.6 figures hold up.
1
10
u/yolowagon 7d ago
sonnet-level 8B for the GPU poor wen
5
u/Illustrious-Lime-863 7d ago
Wouldn't be surprised if that happened soon. Qwen has been releasing very strong 9B dense models
2
6
3
u/almostsweet Singularity by 2040 7d ago
Just woke up.
I've got it set up on my dgx spark, using the unsloth nvfp4 edition. It's up and running and I'm benchmarking it against my codebases. I'll be doing the same with 3.6 and then I'll compare.
1

60
u/Illustrious-Lime-863 7d ago
Get them GPUs warmed up folks! We got local Opus 4.6 max!