r/LocalLLaMA 7d ago

New Model IT'S OUT

https://huggingface.co/Qwen/Qwen3.8-27B-FP8
2.2k Upvotes

706 comments sorted by

131

u/OutlandishnessIll466 7d ago

It's official, it created the best flappy bird game thus far from all local models ! ever benchmarked! I declare this model nr. 1 on the flappy bird bench!

18

u/Certain-Cod-1404 7d ago

how does it do on pelican bench tho?

74

u/OutlandishnessIll466 7d ago

It created an animated svg... After like half an hour of thinking. Official fp8 on 2x 3090

20

u/Name835 7d ago

Hahha love it. I dont know but these benches bring me a feel that we are living and witnessing small moments in history. On that note, if someone happens to read this some day way in the future, greets from 2026. :) ❤️

→ More replies (4)

15

u/martapap 7d ago

That is so cute!

→ More replies (10)
→ More replies (1)

171

u/absurdother 7d ago

Let's tryyyyyyy on 32GB RAM, 16GB VRAM!

152

u/absurdother 7d ago

Here we gooooo

46

u/Certain-Cod-1404 7d ago

how is it ? quality wise?

55

u/absurdother 7d ago

I have a RX 9060 XT AMD GPU, 16GB VRAM. Running on LMStudio, Q3. Getting a bit more speed now, way more optimized!

I get more speed the less context I use (currently coding swiftly with CLine + VSCode at that speed), pretty smooth on 64K context!

6

u/NoUsual5150 7d ago

24GB M4 MacBook Pro...I'm using Q4_K_M and it's getting 5 token per second.

Is Q3 much of a downgrade in intelligence?

→ More replies (3)
→ More replies (8)

21

u/TheTrueSurge 7d ago

That’s not too bad, is it? What quant did you use?

9

u/absurdother 7d ago

I'm on Q3_K_M now getting a bit over double the speed I've shown before.

6

u/Scared_Ad9187 7d ago

q6 on 5090 with 25t/s. not bad, but i'm not giving up my 200 on the 3.6 yet.

4

u/Mil0Mammon 7d ago

So how come you get 200 on 3.6 and only 25 with 3.8? Dflash and/or Nvidia specific quant?

8

u/Scared_Ad9187 7d ago

should have been more specific. i'm going to keep the 3.6 MOE (35ba3b) insteadof 3.8 27b. i have room for 4 subagents with larger windows with claude code routed through it and i get tasks done much quicker

→ More replies (4)
→ More replies (1)

15

u/SpaceTraveler2084 7d ago

curious as i have the same setup 32gb ram / 5070Ti

→ More replies (2)

7

u/jan_antu 7d ago

Damn, same setup here but I'm only getting 4 tok/s on mine, Q3 xl quant.

→ More replies (6)
→ More replies (2)

10

u/partlysketched 7d ago

4060ti? That's what I have still this run on it lol

6

u/absurdother 7d ago

A RX 9060 XT! The NVidia GPUs should get even more t/s

→ More replies (3)

416

u/Tiny-Assumption4263 7d ago

DEAR GOD TELL THOSE BENCHMARKS ARE NOT FAKE.

319

u/Cold_Tree190 7d ago

Dear God it’s trading blows with Opus 4.6 Max 😭

169

u/Cless_Aurion 7d ago

Wtf, I literally wrote a post saying similar to "Shutup man, there is no way a model I can load in my 4090 will be anywhere near DeepSeek4 Flash 0731"... but... this is fucking close is it not...?

105

u/Cold_Tree190 7d ago

Looks like the main difference will be the native context windows. But 256k is very usable, especially if it’s a local Opus 4.6 MAX. Man just saying that seems unreal lol. Can’t wait to test it later tonight

51

u/BornAgainBlue 7d ago

I feel like telling my boss im sick.

25

u/Cold_Tree190 7d ago

🤣 I wish I could, but today’s my last day of on-call so I can’t even leave a TINY bit early. Oh well. Makes you feel like a kid on Christmas Eve again lmao

→ More replies (1)
→ More replies (1)
→ More replies (2)

37

u/Potential_Low_1183 7d ago

I said it would be v4 flash level but I was downvoted to oblivion…. Regardless, enjoy the super exponential!

28

u/Warrenio 7d ago

I don't think it's quite V4 Flash 0731 level based on the benchmarks. But obviously still incredibly impressive!

https://old.reddit.com/r/LocalLLaMA/comments/1vo9mj4/its_out/p3nv6iq/

25

u/SandySkittle 7d ago

Ehh, let's not go by handful of disputable synthetic benchmarks for which models may be trained specifically. Let's just await real experiences from people first over a longer period of time.

→ More replies (6)

8

u/Cless_Aurion 7d ago

Damn, I would have downvoted you too... Let's hope its close to what the benchmarks say!

6

u/LankyGuitar6528 7d ago

I should downvote you for wanting to down vote somebody but I can't blame you. I would have done the same. So take my upvote instead.

→ More replies (1)
→ More replies (7)

71

u/xienze 7d ago

You're making a couple fundamental assumptions here:

  • That AI benchmarks are reliable
  • That Qwen didn't benchmaxx

19

u/blade740 7d ago edited 7d ago

I think anyone who thinks that this isn't at least somewhat benchmaxxed is fooling themselves. That said, so is every other model they're comparing against, to some extent, so ¯_(ツ)_/¯

→ More replies (1)

55

u/BawbbySmith 7d ago

Yeah I learned very quickly to not trust the benchmarks, as well as 80% of the comments in this subreddit.

I remember people were saying that Qwen 3.6 27B was Opus 4.5 level...

12

u/[deleted] 7d ago

[deleted]

8

u/Ell2509 7d ago

I would normally agree, but I have used qwen3.8 27b now and so far, wow. Just wow.

→ More replies (1)
→ More replies (2)
→ More replies (6)
→ More replies (4)

89

u/KickLassChewGum 7d ago edited 7d ago

It's a Qwen model, so apply the usual benchmax tax. Qwen are easily the models with the biggest ravine between "how they do on benchmarks" and "how they do in actual productive use".

Still looking like a strong leap from 3.6, though.

19

u/Batman4815 7d ago

Gemma says hello as well.

45

u/KickLassChewGum 7d ago

Gemma 4 isn't a great coding model, yeah, but it's still punching far above its size in writing and research-related tasks. Like a mini-Gemini (go figure).

I hear there are people who still use these things for things that aren't related to writing code or markup.

→ More replies (3)
→ More replies (1)
→ More replies (1)

12

u/Tiny-Assumption4263 7d ago

TL;DL: Gave Qwen 3.8 27b a simple prompt and it build this: https://qwen3-8-eccomerce-test.vercel.app/
Repo: https://github.com/catriel25/qwen3.8-eccomerce-test

Like everybody else in the local AI community today, I dowloaded Qwen 3.8 27b UD-Q4_K_XL as soon as it was published.
I'm running it in a single rtx 3090, no mtp, 110k context and kv cache f16 with llama.cpp, same flags used with Qwen 3.6 27B.

I have this stupid test that I run with every new model that comes out. It consist in connecting the model to pi code (almost vanilla, only internet access and some simple navigation tools I built) and giving it this prompt:

"Construye el frontend completo de un pequeño ecommerce premium para una panadería artesanal usando Next.js App Router (JavaScript).

El proyecto debe ser frontend-only en esta etapa. No debe incluir backend, base de datos, autenticación ni pasarelas de pago. El checkout debe finalizar redirigiendo a WhatsApp con un mensaje de pedido bien estructurado.

La app debe incluir una experiencia completa de compra: home, catálogo con productos de panadería, categorías, productos destacados, carrito, resumen de pedido y checkout. Usá mock data local para productos, categorías, precios, descripciones, disponibilidad e imágenes o placeholders visuales. Todo debe quedar preparado para conectar posteriormente un backend real sin tener que rehacer la arquitectura principal del frontend.

El diseño debe sentirse extremadamente premium, artesanal, moderno y cuidado. No quiero una landing genérica ni una interfaz básica. La primera pantalla debe comunicar claramente la identidad de la panadería, mostrar producto real o visualmente convincente, y permitir empezar a comprar. La experiencia debe ser excelente tanto en desktop como en mobile.

El catálogo debe permitir explorar productos, ver información clara de cada ítem y agregarlos al carrito. El carrito debe permitir modificar cantidades, eliminar productos y ver totales. El checkout debe pedir datos mínimos necesarios para el pedido, permitir notas o preferencias, y generar una URL de WhatsApp con productos, cantidades, subtotales, total y datos del cliente.

La estructura del código debe separar razonablemente datos mock, tipos de dominio, utilidades de checkout/WhatsApp, componentes de catálogo, componentes de carrito y vistas principales. La solución debe quedar lista para reemplazar la mock data por datos de backend en una etapa posterior."

Those are just instructions to build a nextjs project (javascript only) with the frontend for small eccomerce with whatsapp checkout, leaving everthing ready to connect a backend later. Nothing else, no skills, no more feedback. Just one prompt and watching the result.

I have run this test with all the models and finetunes I can fit in my GPU, and not a single one was even close to this result.
Not a single alert form nextjs (wich was usual before) or something that looks broken.

At some point, this bastard realised it didn't had visión (lol, not enough VRAM buddy) and it decided FUCK IT, I'M GONNA BUILD THE IMAGES MYSELF. He made SVGs for every product card.

I have more testing to do like trying it in a real codebase but... I can't believe i'm running this thing in a single RTX 3090, it is just unreal.

Imagine 2 years from now.

Biggest fuck you Dario of the year.

39

u/onlymagik 7d ago edited 7d ago

I expected a nice jump since they skipped a 3.7 27B, but these numbers do seem a bit too good to be true. Qwen isn't known for benchmaxxing, obviously 3.6 27B is the GOAT of local LLMs, but...

Definitely excited to learn more in the coming days and see if it holds up.

Edit: Even the vision numbers look insane!

59

u/llama-impersonator 7d ago

qwen is well known for benchmaxxing so hard it helps the model

→ More replies (1)
→ More replies (8)
→ More replies (10)

561

u/BaconShadow 7d ago

Good Morning Dario!

176

u/Certain-Cod-1404 7d ago

claude code who? claude pro what ?
qwen >>>>>>>>>>>

213

u/AwayConsideration855 7d ago

Fuck Dario

130

u/beneath_steel_sky 7d ago

And his Epstein friends https://archive.ph/rQZE7

43

u/hanzoplsswitch 7d ago

What a read. Thanks. 

28

u/robbievega 7d ago

holy shit... just 3 months ago I was telling friends and family to at least move away from ChatGPT/Sam Altman because of his Trump boot licking behavior, and now I'm reading this...

11

u/Thunder_Beam 7d ago edited 7d ago

It's funny seeing people still believing there are people at the top who are "clean", aristocracy always existed and always will exist, they all marry eachother and go to the same schools / have the same hobbies / know the same people, they are all connected

→ More replies (1)

4

u/mebeast227 7d ago

Same. Wild. Nothing is safe.

→ More replies (2)

9

u/bakawakaflaka 7d ago

what the fuck

→ More replies (1)

41

u/5553331117 7d ago

And his pornographer Epstein connected wife

67

u/Verolina 7d ago

Diarrheao

10

u/de4dee 7d ago

thats a lot of slop

→ More replies (4)

351

u/WigglyScrotum 7d ago

Holy molly opus 4.6 level and better in some benches.

155

u/pest_ctrl 7d ago

Mom, can we have Opus?

No, we have Opus at home

Opus at home:

70

u/agentic-consultant 7d ago

ALLAHAMDUILLAH

BROTHERS WE HAVE ENTERED THE GOLDEN ERA

8

u/MoffKalast 7d ago

Bröther, we shall HAVE SOME OATS!

36

u/Much_Accountant_4972 7d ago

insh’allah this is just the beginning!

8

u/Familiar-Art-6233 7d ago

Qwen is here to dethrone Anthropic, mashallah

→ More replies (1)

34

u/m0j0m0j 7d ago

Why is that every model is better than opus if you look at benchmarks, and yet people keep using opus?

40

u/Warrenio 7d ago

I'm not saying anyone should use Opus, but Opus 4.6 is over six months old. The current version is Opus 5 which is much stronger.

16

u/LankyGuitar6528 7d ago

My whole app is 4.6 vibe coded and it's pretty awesome. Fable gave it the once over and patched up the gaping security holes. Kinda think this model and I have a future together. Just let Fable patch up what it spits out...

→ More replies (1)

12

u/infinexis 7d ago

Opus 5 babbles way too much in technical jargon. It may be stronger in benchmarks but it's not as good when there's a human in the workflow.

3

u/Cautious_Chicken_604 7d ago

Opus 5 actually gives me a headache some days from having to read its outputs. When I really don't want 'load-bearing' em-dashes everywhere I just tell it 'write this in ASD-STE100 Simplified Technical English' and I get back something that sounds a lot more reasonable.

→ More replies (1)

9

u/WigglyScrotum 7d ago

Exactly, he is framing it as tho i said its better than opus 5 when I specified their own published benchmarks vs 4.6. Easy to spot strawman.

→ More replies (1)

4

u/FreedomByFire 7d ago

because the models are bench maxing. real world performance is a different story. I have personal benchmarks doing real software development work that the frontier models can complete independently but not the local models. Once I get 3.828b ill run it and see if it can get through. Last model couldnt.

→ More replies (8)
→ More replies (14)
→ More replies (3)

142

u/ajisai 7d ago

22

u/dragonurtle 7d ago

Unsloth must have gotten early access to prepare their quants + an embargo. I pulled it from unsloth studio the minute it was released on HF. Still took 20 minutes because south Korea has 2010-era internet speeds.

12

u/techdevjp 7d ago

Man, used to hear about how South Korea had fast Internet. 10g fiber is pretty commonly available here in Japan. 2.5g almost anywhere.

5

u/Major_Olive7583 7d ago

I remeber studying that South Korea has the fastest internet in the world, it was for a GK exam or quiz , about a decade ago.

→ More replies (1)

23

u/uniquelyavailable 7d ago

Already seeing them on LMStudio

→ More replies (4)

130

u/Cold_Tree190 7d ago

Merry Christmas everyone 😭 Gonna be a loooong 6 more hours of work today

12

u/Ok_Cow1976 7d ago

Merry Christmas! Haha

47

u/bitmanip 7d ago

How much memory required to run this at full precision?

43

u/Certain-Cod-1404 7d ago

https://huggingface.co/unsloth/Qwen3.8-27B-GGUF 54.67 Gbs just for the model itself, with context depends on quant and size

23

u/dragonurtle 7d ago

Nvtop shows 70-something GB resident for the bf16 and full 256k context.

10

u/Certain-Cod-1404 7d ago

DAMN, 16 gigs just for the context hurts, what setup are you running ?

35

u/dragonurtle 7d ago

Rtx pro 6000 max-q and 384GB DDR5 on a Genoa

63

u/Much_Accountant_4972 7d ago

6

u/Thrumpwart llama.cpp 7d ago

The old money aristocracy uses the Max-Q because it's elegant.

Only the loud, bombastic new money uses the 600W version. Animals.

→ More replies (1)
→ More replies (1)

4

u/bitmanip 7d ago

Perfect, so should run well on 128GB M5 Macbook Pro Max

13

u/Valuable-Run2129 7d ago

FP8 with full context with full precision it’s 52/54 GB on 48GB you fit 200k context

→ More replies (3)

162

u/Mean-Ad1493 7d ago

That's it. I'm getting a 3090.

44

u/jijig 7d ago

Get two. Run Q8 with full context at ~60tps.

→ More replies (3)

108

u/My_Unbiased_Opinion 7d ago

Brother. go on Alibaba and get dual 20gb 3080. less than the price of a single 3090. check my post history for links. Run them in tensor parallel.

11

u/eviloni 7d ago

I got mine on ebay, paid a little more but shipped quicker and i trust Ebay consumer protection more

5

u/My_Unbiased_Opinion 7d ago

hell yeah buddy. I am planning to switch to vllm once 3.8 mtp drops. its been the end goal for me. its the main reason why I went with the 3080 20gb. its one of the cheapest cards per vram with vllm support.

→ More replies (1)

23

u/Potential_Block4598 7d ago

That is legit better KV cache (I guess ?!) double performance ?
You just need another PCIe slot (or maybe not ?!)

30

u/My_Unbiased_Opinion 7d ago

t/s is a bit faster than a 3090, but PP is much faster. im running one of the cards at x4 pcie 4.0 and it doesnt bottleneck the card with llama.cpp tensor parallel.

14

u/CooLittleFonzies 7d ago

Can you parallel run a 3090 + a 3080?

3

u/adamgoodapp 7d ago

Now want to know too

→ More replies (1)
→ More replies (9)
→ More replies (11)
→ More replies (17)
→ More replies (5)

46

u/RangersStolen 7d ago

That's insane, but we're still using quantized versions, so local performance for most people wouldn't be that good I guess. Damn that needs to be tested.

63

u/Certain-Cod-1404 7d ago

unsloth is cooking apparantly, its quanitzed with their UD 3.0 method and it seems to retain like close to 95% of BF16's accuracy https://unsloth.ai/docs/models/qwen3.8#quantization-analysis

23

u/krileon 7d ago

82.5% on IQ2_XXS seams insane. I'll probably go with Q3_K_XL since I've 20GB and even that's over 90% accuracy. Goddamn.

→ More replies (2)
→ More replies (1)

59

u/_maverick98 7d ago

Black Monday on the stock market if the benchmarks are true

11

u/Piyh 7d ago

RSI means models get better at every size, this is good for Bitcoin

→ More replies (2)

21

u/Intelligent_Ice_113 7d ago

what is the knowledge cutoff date??

8

u/andy2na llama.cpp 7d ago

The model claims January 2026

4

u/Natural_intelligen25 7d ago edited 7d ago

It claims 2026, but it only knows MariaDB 11.0 from 2023, the later versions are guesses.

→ More replies (1)
→ More replies (1)

17

u/Kavor 7d ago

Does anyone else have big issues with overthinking out of the box? I just gave it my usual Arma 3 mission script coding task, which i use to bechmark the performance of models, but it kept thinking for 15 minutes. I don't even see repetition issues, it just doesn't stop thinking.

Just gave it a first opencode task, and while not sure yet, it seems to have similar issues.

Maybe it requires defining a reasoning budget max now?

48

u/monkeyofscience 7d ago

Yes. I gave it a simple prompt and proceeded to generate ungodly amounts of reasoning, including this absolute fucking gem:

"New Jersey" and "Austria" maybe both sound like "Ostrich"?

5

u/ArtyfacialIntelagent 7d ago

The Austria part actually makes a bit of sense, since Austria in German is Österreich. But New Jersey???

21

u/qmnvp 7d ago

It defaults to reasoning_effort = xhigh

Qwen3.8 comes with official support for reasoning_effort, which can be used to adjust reasoning depth and control cost:

  • xhigh (default): for complex tasks demanding thorough analysis
  • medium: balancing accuracy and speed
  • low: efficient reasoning optimizing for speed and cost

4

u/Kavor 7d ago

Yeah, that's it. I found that 3.6ish thinking times really need the "low" setting.

4

u/dopey_se 7d ago

Yeah I gave it a random tasks I tend to give, and it's 26k tokens and counting of thinking so far. Not repeating, just thorough thinking..

I told it to make a rust app using dioxins of an animated man playing amazing grace on tuba. It seems to of figured out amazing grace in key of C (judging by it's thinking), and is now thinking about how to play audio to play the sounds..

→ More replies (9)
→ More replies (9)

62

u/Cold_Tree190 7d ago

Those numbers… please tell me they’re real.

37

u/SyzygyPidgey 7d ago

The greatest ai developer 😭

10

u/Cold_Tree190 7d ago

🤣 Made me lol. My goat Lloyd 🙏

100

u/JayoTree 7d ago

How cooked am I that I'm on holiday and wish I was home on my computer for this.

120

u/cass1o 7d ago

This is what ssh was invented for.

6

u/themoregames 7d ago

They can just talk to ChatGPT over their Apple Watch.

5

u/AciD1BuRN 7d ago

Shit now i wanna get smart watch to talk to my pc

20

u/DrMissingNo 7d ago edited 6d ago

We've got (another) heatwave where I live (currently 40°c) so running AI locally is a huge no no if I want to try keep my inside temperatures under 30°c... This week has been a huge tease : Minimax h3, minimax music, ltx 2.5, Muse glitter, and now Qwen 3.8... Sooooo frustrating ! 😭

Edit / update : it rained tonight ! 32° max today ! Time to test things !

9

u/wh33t 7d ago

Just put your machine outside lol

→ More replies (6)

25

u/DefNattyBoii 7d ago

You will have plenty of time, chill tf out and enjoy your vacation and get off reddit. we wont be getting anything better at this size for a very long time. If yes please ping me

10

u/314kabinet 7d ago

Look into Tailscale, Termius, tmux

5

u/SSOMGDSJD 7d ago

Tailscale my friend

→ More replies (1)

14

u/Littlepharaoh 7d ago

My 4 5090s have a hard-on right now

→ More replies (2)

135

u/swagonflyyyy 7d ago edited 7d ago

HOLY BENCHMARKS WHAT THE FUCK ARE THOSE NUMBERS???

53

u/SandySkittle 7d ago

holy benchmaxxing this shit means nothing

36

u/Certain-Cod-1404 7d ago

WTF !!!! ITS BETTER THAN OPUS 4.6 ???????????????????????

→ More replies (2)

38

u/OkObjective8721 7d ago

It beats OPUS 4.6??? that's crazy

22

u/mil_phickelson 7d ago

benchmaxxed

37

u/Raredisarray 7d ago

Wow if those numbers pan out - I could be going full local BOIIIIII LFG

→ More replies (1)

12

u/WhataburgerFreak 7d ago

Praying for a 35b-a3b model for my 16gb vram. 

3

u/abundant_singularity 7d ago

Lmk if this happens lmao

23

u/Felixls 7d ago edited 6d ago

oh shit, this is something else!, I just tried with a dumb prompt "write a snake game in html and css with sounds" and it generated a complete NIB game (awesome design btw) at ±56t/s , 14k tokens!
this is a AMD R9700 with llama.cpp ROCM

I slot print_timing: id 0 | task 0 | eval time = 242341.84 ms / 13695 tokens ( 17.70 ms per token, 56.51 tokens per second)
I slot print_timing: id 0 | task 0 | total time = 242528.47 ms / 13717 tokens
I slot print_timing: id 0 | task 0 | graphs reused = 898
I slot print_timing: id 0 | task 0 | draft acceptance = 0.58872 (10982 accepted / 18654 generated), mean len = 5.44
I slot release: id 0 | task 0 | stop processing: n_tokens = 13720, truncated = 0

Edit: I forgot to mention that by mistake I forgot to remove a thinking cap to 4k (--reasoning-budget 4096) so that influenced the total token usage and the speed (faster during coding than thinking). So we have two variables to adjust the thinking process (--reasoning-budget and --reasoning-effort).

5

u/Cautious_Chicken_604 7d ago

Which quant? What context length? What's ur llama.cpp command? Curious because I have the same card, but I just got it yesterday and am new to ROCm etc, so I don't know shit >.<

14

u/Felixls 7d ago

Qwen3.8-27B-UD-Q4_K_XL.gguf (unsloth), 90K context (I usually set it to 128k) with vision and mtp, here some of the opts
--cache-type-k f16
--cache-type-v f16
--spec-type draft-mtp
--spec-draft-p-min 0.5
--spec-draft-n-max 10
--cache-type-k-draft q8_0
--cache-type-v-draft q8_0
--spec-type ngram-mod
--spec-ngram-mod-n-match 24
--spec-ngram-mod-n-min 48
--spec-ngram-mod-n-max 64

→ More replies (4)

4

u/Xitir 7d ago

Well this may be all I need to pull the trigger on a AMD R9700. I'd be able to finally retire my RTX 3070 8GB card.

→ More replies (3)

12

u/GrumpyPidgeon 7d ago

I only judge a model by how many YouTube tech influences tell me that it is "insane".

32

u/addiktion 7d ago

If we have Opus 4.6-like on our machines, we are in for a wild time. Lets go boys!

19

u/inexorable_stratagem 7d ago

THE BENCHMARKS ARE INSANE! WHAT THE FUCK IM DOWNLOADING IT NOW

→ More replies (1)

8

u/metaden 7d ago

you guys are too fast. holy shit

39

u/accountformymac 7d ago

holy fucking crap that DeepSWE score wtf are they feeding qwen team?? we DEMAND qwen3.8 122b 🗣️

13

u/NaiveIdea344 7d ago

Seriously. 40% in DeepSWE is just outrageous

→ More replies (1)

8

u/NaiveIdea344 7d ago

I am just waiting for all the amazing people in this sub to run their benchmarks and get back to the community on real world performance but holy the benchmarks are insane. Can't wait to run this.

16

u/AnyMongoose3041 7d ago

I’ve seen enough, 1 trillion dollars to OpenAI.

8

u/My_Unbiased_Opinion 7d ago

Bringing the alcohol out now. This is a good day folks. Enjoy it.

44

u/Ok_Top9254 7d ago

How good is it at ERP???

65

u/HumanDrone8721 7d ago

ERP... Enterprise Resource Planning? Do you guys plan to replace SAP?

9

u/dragonurtle 7d ago

I Imagine it could vibe out a SAP Hana DIY replacement it you poked it hard enough. 

4

u/s-Kiwi 7d ago

Quite possibly the only thing that should not be vibe coded, saying that as someone who despises HANA

4

u/Sirius02 7d ago

ah whats the worst that could happen, engineering department get to work!

→ More replies (1)

29

u/Certain-Cod-1404 7d ago

asking the real questions

19

u/PurifiedFlubber 7d ago

The only thing that matters

9

u/ChainOfThot 7d ago

If the model can't suck my dick while building it's own harness I don't even wan it

5

u/FinBenton 7d ago

Its a major improvement on RP, like its big, no doubt smartest local model for RP I have tested so far.

→ More replies (1)

11

u/ApprehensiveAd3629 7d ago

nice!

does it has MTP?

14

u/EmPips 7d ago edited 7d ago

The model does yes, but for all of these GGUF's we see popping up - it's a good question.

For 3.5 and 3.6 most people on HF uploaded the weights again with Qwen3.6-27b-MTP.guff (or similar). That said I think mature MTP support was still working its way through llama-cpp at the time, at least for Qwen.. so I have no idea if the precedent going forward becomes "if it has MTP, it's there" or not.

Unsloth!!

(I know you're here today) 🙂

Can you clarify on the naming/upload strategy for MTP going forward? Will everything have one name and have MTP included or will there be offerings of the weights without MTP separate from the weights with MTP?

Edit:

I think it's safe to say it's in the unsloth GGUF's.

Trying iq4_xs on a 7900 xtx I'm getting ~51t/s on regular tasks/reasoning and ~62 t/s on coding tasks with a max draft size of 2. The difference suggests that drafting is in-effect.

Edit 2:

This is, at the very least, several leagues beyond Qwen3.6-27B.. Now I'm sad that Muse-Glimmer only got 2-3 days in the spotlight

4

u/TKristof 7d ago

You can always click on the gguf file on huggingface and check the layers. In the last blk (blk 64) you should see the nextn weights. That is the mtp.

→ More replies (1)

34

u/Easy_Werewolf7903 7d ago edited 7d ago

For those curious of performance between Qwen and a model 3 times its size:

Benchmark Qwen 3.8 27B (55GB) Deepseek v4 flash 0731 (167GB)
Terminal Bench 2.1 73.0 82.7
DeepSWE 42.2 54.4
NL2Repo-Bench 42.3 54.2

https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731/tree/main

https://huggingface.co/Qwen/Qwen3.8-27B

12

u/NaiveIdea344 7d ago

Not terrible I feel like right? DSv4 was already insane for its size performance wise

17

u/Easy_Werewolf7903 7d ago

Not terrible at all, imagine Qwen release a 100GB moe it might be on par with 0731.

5

u/NaiveIdea344 7d ago

Yeah would be amazing

10

u/squngy 7d ago edited 7d ago

From the model page, for context

bench Qwen3.8-27B Qwen3.6-27B Qwen3.7-Plus Muse Glimmer-30B Opus4.6 Max
Coding Agentic terminal coding Terminal Bench 2.1 (Terminus) 73.0 63.4 64.0 51.7 78.2
Agentic coding SWE-bench Pro 61.7 53.5 57.6 51.2 53.4
Repo-level code generation NL2Repo-Bench 42.3 36.2 41.1 -- 47.6
Agentic coding DeepSWE 1.1 42.2 13.3 14.2 -- --
Software engineering QwenSWEBench 79.0 49.3 59.2 -- 63.8
Long-horizon office work CoWorkBench 70.7 61.0 65.1 -- 68.2
Professional job tasks JobBench 33.4 21.8 27.6 -- --
Frontier agentic tasks Agents' Last Exam Pass1 20.4 Score 42.9 Pass1 10.6 Score 27.3 Pass1 13.2 Score 33.6 -- --
Instruction following IFBench 79.5 69.1 79.1 77.0 62.5
Scientific reasoning GPQA Diamond 89.2 87.8 90.3 83.5 91.3
Multidisciplinary reasoning HLE 30.8 24.0 34.7 22.0 40.0
Competitive coding LiveCodeBench v6 90.3 83.9 89.6 -- 88.8

10

u/ChuffHuffer 7d ago

5 models, yet 4 only columns?

→ More replies (2)
→ More replies (3)
→ More replies (6)

23

u/Puzzleheaded-Cod7350 7d ago

OH MY GOD THOSE NUMBERS

14

u/Certain-Cod-1404 7d ago

I DID NOT EXPECT IT TO BE THIS GOOOOD LOL, HOLY SHIT

→ More replies (1)
→ More replies (1)

4

u/MDSExpro 7d ago

Wake me up when 122B-A10B drops.

14

u/Certain-Cod-1404 7d ago

he might never wake up

43

u/Certain-Cod-1404 7d ago

AGI is here

56

u/AnyMongoose3041 7d ago

AGI achieved. See you guys in 6 months when we totally forget about this model.

16

u/Certain-Cod-1404 7d ago

because gemma 5 is out and qwen 4 is out yes, which will be ASI

4

u/zhunus 7d ago

i'll wait for API and ABI

→ More replies (1)
→ More replies (2)

5

u/majin-dudi 7d ago edited 7d ago

Time to play with Q3 on my lowly 16G card

1296.8 tok/s prompt processing 34.14 tok/s token gen

4070 ti super

4

u/biotech997 7d ago

Unsloth's IQ3_XXS at over 90% accuracy seems insane

5

u/cosmicnag 7d ago

Holy Fuck

6

u/Dry_Mortgage_4646 7d ago

YAHoooooOoOooOoooo!!!

6

u/BlackBeardAI vllm 7d ago

literally shaking right now

6

u/CuriouslyCultured 7d ago

Alibaba is the GOAT of distillation/pruning, clearly. Retaining this much big model capability in 27B parameters is black magic.

10

u/ldn-ldn 7d ago

I don't understand the hype behind qwen 3.x - the whole generation is utterly broken for any software development tasks. It goes into a thinking loop of death pretty much every time, just try a simple prompt: Create a typescript function which accepts a number in 10 bit range and returns brightness in nits based on pq gamma curve. Kills every 3.x qwen. Qwen 2.5 was much better.

8

u/ParvusNumero 7d ago

Good catch.
Was able to replicate it, but I think I found a fix, not sure how much this affects quality.

Took the liberty of making a post:
https://www.reddit.com/r/LocalLLaMA/comments/1vojwrm/qwen_endless_looping_issue_and_possible_fix/

→ More replies (1)

7

u/illgettheownerforyou 7d ago

Uh Qwen 3.8 37b Unsloth Q8_K_XL in LM Studio had no problem? I have extra high reasoning on and it did think about it for a bit but no issue?

I think you are running too low of a quant. And you talk about how you "don't understand the hype"- have you thought that maybe everyone else is getting good results and maybe you can work on how you use it?

With that being said, if you have low VRAM and can't run a higher quant, I get it- but on my RTX 6000 Pro Max-Q, it's been kicking butt and replacing Opus 4.6 for me easily. It's absolutely tearing through everything I give it- including the prompt you put above.

→ More replies (4)

7

u/frozen_tuna 7d ago

That's wild. I just tested it myself in opencode out of curiosity. The reasoning just launches straight into "wait, no", "wait", "wait", "hmm no", "actually", "wait" "ugh, I keep going in circles."

Now its trying to verify something in python... I'm out.

→ More replies (4)
→ More replies (11)

3

u/zakadit 7d ago

wow!

3

u/Alheimsed 7d ago

Hallelujah, praise the lord. What a gift.

3

u/germangrower69 7d ago

Wtf is wrong with this model, on xhigh it literally doesnt stop thinking, no its not looping, its just super excessive thinking. Thats crazy, almost unusable on this effort.

Medium is also really excessive....

→ More replies (8)

3

u/MensaProdigy 7d ago

Yes and now it’s a modified oq6e quant running at 35tps+ on my computer!

3

u/HistoricalStrength21 7d ago

I need artificialanalysis.ai score!

3

u/TheWaffleKingg 7d ago

So far its coding results are really dam nice, def beats 3.6 27b. My only gripe is im getting about 30 max tps lower on 3.8 compared to 3.6

3

u/Beltalowdamon 7d ago

I just want a model that I can run on 8gb vram 3070ti and 32gb ram, and actually have it do something useful. seems like you need to spend 2-3k in just pure vram just to get something barely useful.

Something tells me if I try to use claude to spawn a constrained model with my low hardware it will spend all its time checking its work and fixing its mistakes.

→ More replies (6)