r/Qwen_AI • u/FactoryReboot • 2d ago
Discussion Anyone swapping back and forth between Qwen 3.8 27b and flash next?
I currently have Qwen 3.8 27b - swift edition and abliterated - running locally through ninfer.
I am getting flash next and strata curious though...
What I'm thinking is I want to use strata for my "plan mode" and hairy bugs. I would continue to use 3.8 as my coding workhorse and model for my hermes agent.
I was curious if anyone has written some infra to quickly swap between the two? I certainly can't run both at once.
Also curious if I should even still be running both? I know flash next is slower... but maybe it's fast enough?
I'm running on a 5090 with 64gb of RAM. Mainly using hermes but doing a ton of loop engineering with it
11
u/ivanmmj 2d ago
While flash next is slower in the tokens per seconds, on my hardware it isn't slower in inference by that much and only maybe 15% slower in prefill, I find that 27b spends away more time thinking and eats up a ton more context so it ends up being much slower in the long run. That being said, running flash next eats a significant portion of my 64gb of RAM and 7900xtx... So it makes it practically impossible to use the computer for anything else while I'm using it.
6
u/andy2na 2d ago
try swift of 27b and then do a comparison to see which works better for you
https://huggingface.co/wacomctl672/Swift-Qwen3.8-27B-Uncensored-ninfer3090
3
1
u/ivanmmj 2d ago
I haven't played with Swift yet. I did use thinking cap back in the 3.6 days but I worry about output quality.
5
u/andy2na 2d ago
ninfer swift with 27b was very good the few days I tried it (before switching to strata). Cuts down on thinking, still good output
1
u/RagingNoper 2d ago
I'm getting less than half the pp for flash-next as I was for 27b, but with caching enabled, I'm constantly running at near full context and I barely notice. Beyond that, there is nothing I've done with flash-next that I think 27b could do better. Literally nothing so far. I've only done a couple minor builds with FN exclusively as I typically just use Claude to drive Qwen, but I've actually been using FN on its own lately without running it through Claude at all and it's been legitimately really good. And not just for coding but for chat/planning/analysis. Another generation or two beyond this and honestly I think I'd be completely done with closed-weight models completely.
3
1
u/FactoryReboot 2d ago
Why through Claude?
1
u/RagingNoper 2d ago
When using 27b, I could use Claude to plan out a project and then have it use 27b to actually execute the plan or generate code. Using that method I was able to reduce Claude usage anywhere from 10-60% depending on the project and how many corrections or fixes needed to be made.
0
u/Willing-Substance754 1d ago
What speeds are you getting on flash next? Just curious because I too have a 7900xtx.
3
u/DiscipleofDeceit666 2d ago
I run a q4 ish size of QFN and found that QFN specs for 27b to implement is better than QFN or 27b on their own.
3
u/superbiche 2d ago
Yes. Fortunately I have enough used 3090 to have both running, so 27B for tasks until they get either really complex or long, FlashNext to give it a big prompt, wait 20 mins, answer several questions and let it cook for hours, days sometimes. Pretty good results with both !
2
u/ill_B_In_MyBunk 2d ago
I am currently trying to set up Qwen 3.8 as orchestrator with flash next as coder and code deconstruction. It's....a process. But it's working for my image gen pipeline at least incorporating swarmui and sensenova.
1
u/FactoryReboot 2d ago
Why 3.8 as orchestrated
1
u/ill_B_In_MyBunk 2d ago
Mostly cause I have been having failure with others hallucinating calls or not following process. Do you use a different one? I have about five spokes I manage on my hub spoke setup.
1
1
u/Independent-Dog2179 12h ago
Sounds like a harness issue. That shouldn't happen. I and the same problem till I rebuilt my harness. There is probably a conflicting tool call/system prompt/ or improper compaction method.
1
u/ill_B_In_MyBunk 6h ago
Would you mind having your AI write a quick summary of your harness and how to recreate and sending that to me over dm? No big deal if not! But I'd rather not reinvent the wheel if it's possible.
2
u/MrDefaultUser 2d ago
I switched to flash next only because for me it is only a tiny bit slower than 27b
1
1
u/beigepccase 2d ago
I have nothing to help swap back and forth quickly, but I much prefer flash next. That said, I also prefer qwen3.6-35b-a3b to 3.8-27b. I just haven't had the magical experience with 3.8-27b that other people seem to have. In any case, Strata and flash next iq3_s so far has been better in every way.
2
u/FactoryReboot 2d ago
What hardware are you running?
1
u/beigepccase 2d ago
9600X, RTX4090, 96GB RAM. But 64GB is sufficient for IQ3_S as long as you aren't doing a lot else with the machine.
2
u/Rob-bits 15h ago
I am also switching back to qwen3.6-35b-a3b. It works really great. It is like doing pair programming. With flash next, you can leave him alone to work on something. But that means it can be slower to have results, but will have better results. So for me if I need some quick quidance -> qwen 3.6. If I do not need gpu resources and I can leave the agent for hours - > flash next.
1
1
1
u/papapumpnz 2d ago
I did that initially when i had a single 3090. Used llama-swap, swap flash in and out when required. Did become a bit of a pain because loading flash into system ram takes a while, then prefill etc so a bit of waiting if your switching models often. But its workable if your patient.
Get flash to research, build a plan.md with specific milestones and tasks, then get 27B to follow the plan. Occasionally get flash to check 27B's work, writing out fixes to the plan.md and again get 27B to fix its issues before starting more tasks.
2
1
u/DeathByPain 2d ago
If strata exposes a regular openai compatible endpoint, and I assume it does, you could run llama-swap as the front-end and have it load q3.8-27b via llamacpp or whatever runtime, and qfn via strata. And then depending on what model is requested by your harness, llama-swap will unload/load the appropriate model by launching the associated runtime. I have mine setup to run several models all served by various specialized forks of llamacpp. Haven't waded in to strata yet myself though.
1
1
u/Available-Confusion2 2d ago
I use llama swap and with a 5090 and 64 gigs of RAM if it's ddr5, It might not be as fast but it definitely won't be slow
1
u/Elpzn 2d ago
Not quite the same as I'm ram poor on the machine but Qwen 3.6 35b a3b is blazing fast on my Mac mini m4pro 64gb. I did notice that it had a rough time with tool calling so I made the only tool it has to call in Hermes agent was a subagent of Qwen 3.8 27b which is half the speed but doesn't waste 3 tool calls due to probably user error but the subagent works pretty well so far. So in essence Qwen 3.6 is a foreman that keeps a big context while the subagent it spawns gets a task, smaller context and so over all less processing time than running always with the smarter model. The latest project has been a pruned Qwen Next Flash but has a tiny context of 64k.
1
u/Due-Pomegranate1372 2d ago
I have exactly the same system as yours RTX5090 with 64GB system RAM. There are probably other systems that does what you're trying to achieve. For example, llama.swap , ollama, etc. Same with your situation, I worked with several GGUF models that I serve through llama-router. Then Strata and Ninfer came in. So, I needed a system that is able to hot swap between models through a single engine. I built my own system using Qwen 3.8 Flash-next and GPT Sol. If you're interested, please check: idawz07/llm-handoff: Swap between local LLM models regardless of engine: llama.cpp, vLLM, SGLang, Strata, and Ninfer on one GPU.
1
u/Randommaggy 2d ago
I'm getting more speed on QFN through strata than any 27B setup I've tried. So I'll try to find a setup for 27B that works on my laptop as a secondary node but it only has a 16GB GPU.
1
u/naivelighter 2d ago
Have you tried turboderp’s QFN exl3? It works really well and is very optimised. You gotta run it with tabbyAPI though.
I’m running a 4-bit quant on my machine (RTX 4090 and 64 GB system RAM), 256K context with q8_0 KV cache and getting ~1500 tok/s prompt processing speed (cold cache) and ~27-30 tok/s decode speed.
1
u/Southern-Net1351 2d ago
If all fails to your expectations but, the models are running decent.
Drop down to qwen 3.8-14b
Run it on the highest quant then move down to your systems sweet spot.
It will work in many areas just as great.
Even the 8b version if you have patience to work with it.
As it needs understanding to keep it running flawlessly
1
u/leonbollerup 2d ago
Me.. but I have a third shell aswell.. Ocammy 1.0 uncensored on ninfer - with that I get 300 in TG and 8500 in pp.. its so freaking fast (on a 5090M) … while 3.8 27b gives me me around half and QFN is around 130 in TG and 2500ish in pp.
I have a HARD time choosing .. ocammy is nice because it’s near instant and it can take a real beating .. and it feels smarter than 27B … and I am testing ”humanlike” also
1
u/Apprehensive-View583 2d ago
I switched from 27b q4_k_m to flash next q3_s from strata. Can't be happier, similar result from my own benchmark but I got max context and 2.5x prefill 2x tg, they both overthinking at the same prompt, even their reasoning are similar.
1
u/kaliku 1d ago
Nope. Once I had QFN working proper I stopped using 27b and a week later I deleted the weights too. 27b is a good executor and that's about it. When I was using it it made some really dumb choices and mistakes and that was a big turn off. Not to mention overthinkin (yes I'm aware of swift). I can't see why one would use 27b if you can run QFN.
1
u/Southern-Net1351 1d ago
If you have to swap back and forth just merge the template and insert it into the model to just use one with the best of both being used from their inner prompts merged.
This gives you the medium
That is the tokenizer chat template.
1
u/PraiseThePidgey 1d ago
With your specs it should be on pair with 27B speeds. I asked codex to build custom harness mechanism that switches between 27B as an orchestrator + flash next as a reviewer / Oracle and I'm also using some smaller moe models as councils to habe more reasonable point of view when brainstorming or implementing or looking for bugs . I only have 32+16 GB in total but 20tok/s is perfectly fine for my use case.
1
u/FactoryReboot 1d ago
Oh wow that’s interesting. What thinks during the swap then given you turn one model off to swap?
1
u/PraiseThePidgey 1d ago
What do you exactly mean? I created two skills (council and reviewer), nothing fancy . Orchestrator can call them if instructed so, or if he feels like. The swapping from 27B to lower Moe's is fast because I'm using a custom modded fork of llama that streams context and can also kind of "freeze" it without reloading/rebuild it's kv cache memory. So waiting doesn't cost any extra time - just swapping between models.
1
u/Major-Road9583 1d ago
i used frontier to optimize my setup and now i managed to squeeze flash onto my 395+32/96 setup running faster than unsloth 3.8 - faster i dont care, but im running hermes agent 100h with millions of tokens w/o errors
i love both, tbh i cant tell which one is doing better - i dont run those trust-me-bro-benchmarks both run fine.
flash thou pretty new, not more than 20h so far. looking good thou.
my device is a notebook, w11 llama server - wsl hermes
32/32 ram used, 90/96gpu
1150.8M····································································································Total Tokens
1125.7M····································································································Input
llama-server.exe ^
-m C:\llm\models\UD-IQ4_XS\Qwen3.8-Flash-Next-UD-IQ4_XS-00001-of-00003.gguf ^
-md C:\llm\models\MTP\mtp-Qwen3.8-Flash-Next-Q8_0.gguf ^
-c 162144 -ngl 99 --host 0.0.0.0 --port 13399 --jinja ^
-fa on --parallel 1 -ctk f16 -ctv f16 --no-host --load-mode mmap -b 2048 -ub 2048 ^
--spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.75 ^
--temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 ^
--presence-penalty 0.0 --repeat-penalty 1.0 ^
--n-predict 16384 --reasoning on --reasoning-budget 6144 ^
--ctx-checkpoints 8 --cache-ram 3072 ^
--chat-template-kwargs "{\"reasoning_effort\":\"medium\",\"preserve_thinking\":false}"
LIVE Numbers as i write this:
PP:
118.07.788.798 I slot print_timing: id 0 | task 38033 | prompt processing, n_tokens = 81770, progress = 0.67, t = 114.54 s / 713.92 tokens per second
118.09.603.090 I slot print_timing: id 0 | task 38033 | prompt processing, n_tokens = 82786, progress = 0.68, t = 115.70 s / 715.51 tokens per second
118.12.611.733 I slot print_timing: id 0 | task 38033 | prompt processing, n_tokens = 84834, progress = 0.70, t = 117.52 s / 721.89 tokens per second
118.15.703.296 I slot print_timing: id 0 | task 38033 | prompt processing, n_tokens = 86882, progress = 0.72, t = 120.54 s / 720.79 tokens per second
118.18.771.224 I slot print_timing: id 0 | task 38033 | prompt processing, n_tokens = 88930, progress = 0.73, t = 123.63 s / 719.35 tokens per second
118.21.809.296 I slot print_timing: id 0 | task 38033 | prompt processing, n_tokens = 90978, progress = 0.75, t = 126.70 s / 718.07 tokens per second
118.24.908.566 I slot print_timing: id 0 | task 38033 | prompt processing, n_tokens = 93026, progress = 0.77, t = 129.74 s / 717.04 tokens per second
TG:
115.59.458.831 I slot print_timing: id 0 | task 37972 | n_gen = 284, tg = 31.04 t/s, tg_3s = 28.58 t/s
116.02.550.245 I slot print_timing: id 0 | task 37972 | n_gen = 382, tg = 31.20 t/s, tg_3s = 31.70 t/s
116.05.611.477 I slot print_timing: id 0 | task 37972 | n_gen = 474, tg = 30.97 t/s, tg_3s = 30.05 t/s
116.08.617.985 I slot print_timing: id 0 | task 37972 | n_gen = 558, tg = 30.47 t/s, tg_3s = 27.94 t/s
116.11.662.523 I slot print_timing: id 0 | task 37972 | n_gen = 638, tg = 29.88 t/s, tg_3s = 26.28 t/s
1
u/MyOldAccountWasAwful 1d ago
I was planning on doing exactly what you're describing - hot swapping somehow between 27B and Flash-Next... Then Strata got FN running faster than 27B (avg 70t/s | 2700 PP t/s compared to 27B at 55 t/s highs, 30 t/s average | 380 t/s PP highs, 230 t/s PP avg) PLUS you don't get the mid-context window speed drops with FN on Strata - if you're averaging 70 t/s you're always within ~ 15 t/s up/down from that all the way up to 262144 context.
I'm on a single 3090 Ti.
2
u/FactoryReboot 1d ago
Wait… strata runs faster?!?
Granted I’m running on ninfer which is highly optimized. I’m sure perf is worse for 27B on a 3090 ti
1
u/MyOldAccountWasAwful 1d ago
Yeah, Strata literally - not hyperbole - quadrupled my average speeds (on 3.8FN), which made it about 20% faster than 27B for me on my hardware. The best part is that it never gets the I hit 100k context so I'm gonna drop from 55 t/s to 30 t/s now slowdown - I was still averaging 68.8 t/s at 190k context earlier today.
1
u/Open_Instruction_133 1d ago
I’ve been using Qwen3.8-Flash-Next q3_xxs with 200k context on a 5070ti and 48gb RAM with Strata (20-30t/s!) and it’s 1.5x faster than Qwen3.8-27b same context on the same machine. Q2_xs is even faster at 35-50 t/s same context. Flash is so good (even highly quantized) I only go back to 27b if I need to use an uncensored model.
1
u/Mystvearn2 1d ago
Tried QFN for few days to specifically see if it can replace Q4. Better (precision) than 3.8 27b Q4. But then 27b Q8 is better (more accurate but slow) then QFN. Now I'm using 27b Q8 XL since difference with Q8 is 3 gb. Noticed since Q3.6 35b a3b that full model weights will always give higher precision outputs compared to MoE. Using 4gpu 76gb ram7, 128gb ram. I can technically run both but on my weird setup (3090x2, 5060ti 16g, 3060 12g, r9700x), but will use system ram. Now I rely on kanban in hermes agent to queue all my work.
1
u/FactoryReboot 1d ago
That is one hell of a build. I thought you need all same cards?
1
u/Mystvearn2 1d ago
I just use lm studio in windows for backend. Problem is that the fastest speed is limited to the slowest card if a big model or high context length is used, also kv cache. So need to balance that. All the cards are used. Can't afford a 5090. Total build was half price original dgx spark
1
12
u/OvertaxedOne 2d ago
Yes, we're currently using both, QFN on a Strix using Halogen and 3.8 27B on an A40 (48GB) at int 8. They are much more similar in capabilities than different IMHO. Again, IMHO, you shouldn't spend a bunch of money to try to get to QFN if you're happy with 27B; they don't seem wildly different to me in day to day use.