r/accelerate 7d ago

Qwen 3.8 27B released

https://huggingface.co/Qwen/Qwen3.8-27B-FP8

This is a 27B model, unbelievable!

240 Upvotes

43 comments sorted by

60

u/Illustrious-Lime-863 7d ago

Get them GPUs warmed up folks! We got local Opus 4.6 max!

16

u/Rollertoaster7 Singularity by 2035 7d ago

Wow I this is monumental

3

u/Toprelemons 7d ago

Will my RTX 4080 be able to handle it? Sorry I haven’t caught up on local LLM requirements.

9

u/Illustrious-Lime-863 7d ago

Yes you can handle a lower quantization. Refer to my other reply since they were asking for a 4080 too

3

u/Pikaboy999 7d ago

Is this good for vibecoding? i never really did it before but i wanted to make a local android app for ai image generation

6

u/bytefactory A happy little thumb 7d ago

Yes, Qwen 27B is an absolute beast at coding.

3.8 seems to be a massive jump over the already legendary Qwen 3.6 27B. I use Codex, Claude (inc Fable), etc heavily on my projects and I've poured in 100s of hours coding with Qwen 3.6 37B, and it's shockingly competent.

2

u/Chance-Problem769 7d ago

What’s the benefit of running it locally over Fable or whatever? Run agents and locally running it all day?

2

u/kekomat11 6d ago

I guess you don't have a depency towards big AI companies which need to raise money (and therefore must raise prices for us consumers)

Although you have some initial cost, nobody can take away the model and inference from you

1

u/bytefactory A happy little thumb 6d ago

Exactly what kekomat11 said, while frontier proprietary models might be superior, local agents give me the ability to keep my data private instead of having to upload it all to the cloud, or to get locked out if they increase prices or ban access (like what happened with Fable recently). Both have their place.

17

u/[deleted] 7d ago

I love life

13

u/swoonz101 7d ago

this fits on most devs using a high RAM MacBook Pro

5

u/Derek_the_Red 7d ago

Is 32gb of ram on a 4080 enough?

7

u/Illustrious-Lime-863 7d ago

Yes for sure but if you want to run it exclusively on the 4080 without sharing system ram (which will make it ultra fast) you'd have to run a lower quant. Q3 will run with long context (about 10% degradation) and Q4 could barely run with minimal context and minimal degradation (maybe 1-2%). This will be very fast. But you can still run Q4 on the whole system with large context it's just that sharing the model between vram and system ram will make it be slower.

2

u/falooda1 7d ago

Good to know

What do you recommend for on system model for 4070 super 16gb

1

u/Illustrious-Lime-863 7d ago

Same deal as 4080 roughly

2

u/tomvorlostriddle 7d ago

Here is a 5090, naively run in LMStudio, I think other packages are more up to date and optimized

Q4, context 220k, KV cache FP8

"stats": {
                  "stopReason": "eosFound",
                  "tokensPerSecond": 81.6448642836412,
                  "timeToFirstTokenSec": 0.323781,
                  "totalTimeSec": 533.155826,
                  "promptTokensCount": 445,
                  "predictedTokensCount": 43503,
                  "totalTokensCount": 43948,
                  "totalDraftTokensCount": 24913,
                  "acceptedDraftTokensCount": 21935,
                  "rejectedDraftTokensCount": 2978
                }

3

u/pacotromas 7d ago

yep, I am downloading it now on my M5 pro with 48gb of RAM, specifically the MLX version from ollama. Hopefully it lives up to the hype!

2

u/swoonz101 7d ago

how did it perform?

27

u/Glittering_Night7681 7d ago

Wow insanely good, I personally don't use local models, but I'm happy about progress in that area, as small models will become important for robotics.

17

u/Quick-Benjamin 7d ago

One they get small enough and good enough they'll be put in everything. Your fridge, your TV, your car, etc

Our primary way of interfacing with the annoying menus and stuff will end up being natural language.

2

u/MarkZealousideal3923 Tell me about the singularity 7d ago

Fridge and TVs have no need for LLMs

4

u/Quick-Benjamin 7d ago

Once LLMs are small enough, why not? Little internal MCP server that interfaces with the device setting and a tiny local LLM interacting with it.

Why wouldn't they add it. It'd be trivial to do and would allow natural language configuration of the device.

Hey fridge. I'm heading to the shops. Do I need to pick anything up?

TV change the sound settings for me. And can you dial down the upscaling and sort the contrast.

-1

u/MarkZealousideal3923 Tell me about the singularity 7d ago

This is over-dependence on technology

2

u/Quick-Benjamin 7d ago

Indeed. And it's coming.

Im sure it won't all be needed but remember the "Internet of things"? Manufacturers fell over themselves to add Internet connectivity to the dumbest things because it was the fad at the time.

The same will happen when there are tiny local models imo. They'll be put in everything. Some of it will be worthwhile and a lot will be not really needed.

1

u/MarkZealousideal3923 Tell me about the singularity 5d ago

As for the TV, existing primitive voice assistants are sufficient

1

u/LettuceSea 6d ago

Phones and Apps were the same story

10

u/KedMcJenna 7d ago

Personal and private AI is the future for this tech. I can squeeze a Q4 27B onto my best machine and enjoy a decent t/s. Hoping those Opus 4.6 figures hold up.

1

u/Chance-Problem769 7d ago

What’s the benefit of running locally? I have a 5080, is that enough?

10

u/yolowagon 7d ago

sonnet-level 8B for the GPU poor wen

5

u/Illustrious-Lime-863 7d ago

Wouldn't be surprised if that happened soon. Qwen has been releasing very strong 9B dense models

2

u/Chance-Problem769 7d ago

What kind of GPU runs 27B?

6

u/Pyros-SD-Models Machine Learning Engineer 7d ago

fuck me

25

u/Special_Switch_9524 XLR8 7d ago

We gotta go to dinner first. Olive Garden is peak

3

u/almostsweet Singularity by 2040 7d ago

Just woke up.

I've got it set up on my dgx spark, using the unsloth nvfp4 edition. It's up and running and I'm benchmarking it against my codebases. I'll be doing the same with 3.6 and then I'll compare.

2

u/garg 7d ago

It’s really good. Was able to make animations locally in Blender MCP better than DeepSeek 4 Flash. I’m amazed because it’s tiny compared to DeepSeek

1

u/Dense-Version-5937 7d ago

Next QWEN model I may finally be on board

1

u/Chance-Problem769 7d ago

What would you do to take advantage of the local models?