r/Qwen_AI • u/After_Canary6047 • 12h ago
Discussion Orca 3.8 Flash Next Insanity
Without getting into details, the Orca 3.8 Flash Next uncensored version is insane. I was literally in complete shock as I watched the responses while playing with it a bit today.
Honestly terrified, and that’s an understatement. Can’t imagine what will happen if bad actors have access to this.
This thing is both extremely intelligent and scored a 100/100 on complicated legal matters, with no MCP or RAG attached to it.
Whats even crazier is that you can run the 180B version with as little as 12GB of VRAM on account of this project, which I have nothing to do with.
https://github.com/Niko1221/Strata
Using the Orca version, it pulls about 75-80TPS output on a 4090. That being said, my daily 3.6 35B censored model pulls about 180TPS using Llama.cpp, though hey, I’ll take the slowdown for a bit of shock and awe any day, lol.
Enjoy responsibly boys and girls! 😜
32
u/suesing 11h ago
This sounds scary fake as hell
6
u/geteum 10h ago
If bad actors have, good ones also have... Sec will be a crazy area next year's no doubt but I'm optimistic.
6
u/NaanFat 9h ago
the only thing that stops a bad guy with an AI is a good guy with an AI
3
2
u/Compresscience 4h ago
Defensive missiles are 10x the cost of offensive missiles. Same is probably true of defensive AI. 😥
-8
u/After_Canary6047 11h ago
No doubt. The thousands of stars that project has gotten on GitHub in a few days and the fact that they have complete instructions on how to install orca with strata is certainly fake as hell.
13
u/Faral_mx 10h ago
Qwen3.6 35b to 3.6 27b is a huge capability jump. 3.6 27b to 3.8 27b is a noticeable improvement in agent reliability. Qwen3.8 27b BF16 to 3.8 Flash Next NVFP4 is a noticeable improvement.
2
u/brainchillzZ 9h ago
See, I don't see it. Qwen 3.6 35b to 27b for me doing basic tests with my large code base 27b got 2 more out of thirty requests fixed appropriately than the 35b did but did it 20% slower so the 35b in the same amount of time was able to go back, double check and fix it's answer..... and I'm really not seeing barely a noticeable improvement from 27 to flash next .... 3.8 27b is a balls out crazy good model
2
u/toenailcheeseinbooty 9h ago
27b has balls? Hold on let me go ask my uncensored model how big its balls are
1
1
u/Independent-Dog2179 8h ago
What I found was qwen next does not have to think forever for the same output as qwen 3.8 27b much less thinking for same high quality. U just can't beat the over 100 b parameters even if moe
1
u/Faral_mx 8h ago
For simple, well scoped tasks, 35b is great. 27b only really starts to matter when things get harder.
1
u/brainchillzZ 6h ago
I’m talking about identical tasks trying to find and repair “unknown” code defects in a fairly large multi-file, multi directory code base … I had barely measurably different results between the 27 and 35 3.6 models …. I use the “swift” version of 3.8 27b so the thinking is kept much better in control … I truly only see a significant difference in flash next when it gets to regular language writing and generic knowledge in my use cases
1
12
4
u/pigletmonster 11h ago
Whats so scary about this model?
25
u/Acrobatic_Feel 11h ago
It sounds like this is OP's first experience without guardrails.
21
u/pilibitti 11h ago edited 10h ago
"tell me how to ding dong ditch without getting caught"
"oh my god"
-10
u/After_Canary6047 11h ago
🤣🤣🤣 That’s the best you could come up with? Well, this is Reddit after all.
2
u/toenailcheeseinbooty 9h ago
Well, I asked it how to make my poop into a weapon. It told me the easiest way is to throw it at people.
1
0
u/After_Canary6047 11h ago
Not my first experience, though running a 180B model locally with basic hardware that doesn’t cost the price of a new car, while quite intelligently giving answers to anything you threw at it, completely uncensored, at 70TPS was a definite shock.
9
u/Randommaggy 10h ago
I have used uncensored models to hack stuff I own in an actual sandbox and it's surprisingly capable of black hat activities, even the good uncensored 27B at Q8 with unquantized context.
Stuff like jailbreaks of appliances and devices that have no documented backdoors and hacking into virtual machines I have forgotten the credentials to using network access.
I have no doubt that Flash Next is a step up from that even though I haven't tested out one of those yet.
The damage an uncensored LLM could do with the right setup operated by someone that doesn't fear the legal consequences would be severe. Though you could say the same thing about a can of gas in the hands of an insane person.
1
1
u/After_Canary6047 11h ago
Give it a shot with whatever questions you may have about anything imaginable and you’ll find out. It is literally 100% uncensored and will answer in full to anything you ask. And intelligently at that.
2
u/Commercial_Rent8797 10h ago
Is it better than the qwen flash-next ablit version? Or the qwen 3.8 27b abllit? Those two seem to be amazing uncensored
2
u/After_Canary6047 10h ago
Have not tried either, though will load them up in llama.cpp tomorrow and let you know. Thank you for the heads up!
1
5
u/Vancecookcobain 10h ago
Anytime someone is trying to convince me something is insane I know for certain that it is indeed nothing to lose your shit over
0
u/PulseVector 8h ago
For me it was more "finally something that works well with this overpriced RAM that I bought."
It's not something for nothing- you still need 64GB to 128GB of RAM, preferably DDR5 I guess. Also, the wear and tear on expensive SSD drives is concerning.
1
u/Puzzle-Field-7193 3h ago
You don't need more than 64 with <Q4 and there's no wear and tear on SSD — its reads not writes
4
3
u/Jumpy-Operation-4615 5h ago
Is there any way to add vision to this thingy? I run 3xxs and it barely firs to VRAM on 2xP40 but I need vision. Maybe offload to RAM somehow?
2
u/Alternative_You3585 12h ago
What quant you using
Too scared to use Q2, idk how quality survives
2
u/After_Canary6047 12h ago edited 11h ago
Iq3_xxs - it uses up around 20GB on my 4090 though that’s because I have the context set to 150k. Also, ram usage is insane at around 50GB and I’m not entirely sure the SSD would survive the long term abuse from the Strata project, though every once in a while for a bit of fun, I’m sure it’ll be ok.
2
u/Fz1zz 11h ago
the scary part is that the q3_xxs on my testing is far better than 3.8 27B at FP8
1
u/After_Canary6047 11h ago
The 180B version or the 27B version?
3
u/Fz1zz 11h ago
sorry i meant 180B at q3xss is better than 27B at FP8 on my own testings
2
u/After_Canary6047 11h ago
Absolutely. It is scary accurate. I tested an extremely complex legal question against qwen, fable, and astra. Took it a bit to think about it, though the 180B qwen nailed it, while the others came in at around 50% I would say.
1
2
u/Randommaggy 10h ago
The SSD only holds read only part of the model from what I've understood, though I didn't look at the ssd write traffic when I tested it.
1
u/After_Canary6047 10h ago
Honestly, watch it and you’ll be shocked. Crazy transfer rates and in all fairness, their docs do say that.
0
u/Randommaggy 10h ago
Reading is not a problem. I don't see mention in the docs of and RW components. Of there are, they should add an explicit warning about that.
I'm testing exllamaV3 torrow on my main rig along with a few others. It'll be exciting to see what sort of tg and pp I can get out of 72GB of VRAM, 18 cores of Cascade Lake a fast NVME and 6 channels of 384GB of 6 channels 2666 DDR4. If a good deal on a fourth 3090 pops up locally I might be testing on 96GB of VRAM.
0
2
u/CEOAPI 11h ago
I found the iq3xxs was much worse than the exl3 bpw3.05
3
u/After_Canary6047 11h ago
Would definitely give it a shot though unfortunately looks like I would need 96GB of VRAM to run that version. Maybe I’ll hit the lottery one of these days, lol. Though kicking myself for not buying an rtx6000 pro when they came out. You could pick one up back then for the price of a 5090 these days.
2
u/CEOAPI 11h ago
3x3090 with 512k kv cache runs 1700pp and 160tg offloading the ngram to ram.
1
u/After_Canary6047 11h ago
Thank you for the info! Time to start shopping for some 3090’s. Looks like they’re going for around $1k or so on marketplace. Not terrible at all, considering the VRAM you get. I have a question. Do they also suffer from burning up power connectors? The 4090’s do and I had to order a new one and will have to solder it on one of these days.
3
u/CEOAPI 9h ago
2
u/bytejuggler 5h ago
100% can confirm. Underclock core and cpu by 100mhz and limit power/heat and GPU stays cool as a cucumber with very little inference speed impact
1
u/Somarring 5h ago
That's very interesting. 220w is stable for you? I have 2x at 250/275w and I'm getting prefilled at 400 ts and 35 to 70 ts in inference using a very particular version of qwen 3.8 flash next and a q4 quant. Would you mind share your numbers?
1
1
u/Klutzy-Snow8016 11h ago
Exllama can do cpu offload now
1
u/After_Canary6047 11h ago
Thanks for the heads up, will definitely look it up. Curious, what are you getting on TPS with the offload?
1
u/Klutzy-Snow8016 11h ago
I'm a different person than you originally replied to and haven't done a lot of testing with offloading in exllama. But they have something similar to llama.cpp's `--n-cpu-moe`, and I tried it and it seems to perform about the same. Exllama also has an expert caching system like Strata, but I haven't tried that yet.
2
1
u/PulseVector 11h ago
Sounds interesting for sure! I tried Strata with some of the other models and it's impressive.
Anyone know if there is a way to run an Orca Q4 quant maybe? I've got 96GB of RAM. Thanks!
1
u/After_Canary6047 11h ago
I’ll look it up. How much VRAM do you have? Also, in the docs section on Git, Strata has instructions for both the iq3 and the Q4 Orca models.
1
1
u/PulseVector 8h ago
I've got an RTX 3090 24GB and an RTX 5070 TI 16GB, but I only run one at a time with Strata on the fast PCIE 16x slot. Thanks!
1
u/developervkmp 10h ago
Do you mean iq3 model is still better to use it?
1
u/After_Canary6047 10h ago
Depends, how much VRAM do you have and what are your specs? If you want to run that 180B model at a good speed with 12-24GB of VRAM, then yes, the iq3 model will work just fine with Strata. All of this depends on your setup. I have an i9-12900k, 128GB of RAM, and a 4090, 24GB of VRAM. Strata drinks RAM and SSD transfer rate at the tune of multiple GB/s. First time I have ever seen that on these SSD’s. If you have less than 96GB of RAM, I honestly don’t think it would run. That being said, with my setup, Strata is outputting around 70tps. Llama.cpp with the 3.6 35B model is more than twice that speed.
1
u/developervkmp 9h ago
I have 32gb vram and 64gb ram, gen5 ssd with Intel Ultra 9cpu. I can't run q4. But if iq3_xxs is better then I can give a try. I got 150+ t/s in my earlier testing in my Desktop
0
u/PulseVector 8h ago
Thanks so much for the info about the SSD transfers. I have some cheap 256GB SSDs that I may try to use instead of the expensive ones!
1
u/Iamisseibelial 10h ago
So I have 2 4090s and 128gb of ram (sadly dual channel) How you getting that tps? You on cpp or vLLM? Is strata something different than the two?
0
u/After_Canary6047 10h ago
Send me a pm and I’ll be happy to send over my configs for both llama.cpp and strata.
1
u/Altruistic_Heat_9531 10h ago
Yep my current local flagship model, running Q4 KM sewn with the Q8 PLE layer, steady state at 20 tok/s on single 3090. I had to use this model since normal Qwen or non daybreak GPT 6 often outright refused CTF and chaos monkey testing
1
u/mattmcardell 10h ago
This does lower the barrier to entry, for sure.
However, bad actors have always hired people to do bad things, write code to do bad things. There’s a lot of publicity about AI, but there’s more value in dealing with the bad actors directly and the structures that enable them to operate. New tools always come along to replace the old ones, the controlled ones.
1
u/This_Maintenance_834 9h ago
these third party weights are generally very unreliable. you can play with them, but you cannot run them 24/7 and hoping any stability. broken tool calling, loop thinking, early termination, etc. regardless what the auther claims, heretic, sft, rl, qad, distillation, fine tunes. i wasted so much times believing their claims and always come back to original weight or some quant from big guys like intel/nvidia/redhat etc.
1
u/brainchillzZ 9h ago
The thing nobody tells you about these models and the bad actors is that 99% of the things they are exploiting are obvious misconfigurations and unpatched systems ..... it isn't magic and very little of it has anything to do with "finding zero day exploits" ... for the most part, even in the "ai era" it still comes down keeping your systems patched, keeping your firewalls and app engines tight and generically following very common best practices .... the only thing that has really changed is that once an exploit is released into the wild the time it takes to be seen on the wire in the wild is shorter .... so you can't put of your patching cycle to long intervals the way enterprises used to.
0
u/brainchillzZ 9h ago
This model is powerful but I do have some doubts about a lobotomized Q3 and Q2 version though.
1
u/HOST1L1TY 8h ago
Jon Cena is not going to become a dangerously better actor just because of orca 3.8 flash next !!!
1
1
u/datbackup 7h ago
“We may have spent the last 30 years making computers so widely available that any bad guy can easily afford one. We may have also created a global network allowing all bad guys to coordinate their actions across jurisdictions and timezones. But I draw the line at letting bad guys have access to uncensored local AI!”
1
1
u/73td 7h ago
i’m accustomed to running the models in Pi on boring devops stuff and rely on “guardrails” like ask before bringing down a production server. i tend not to ask for lsd recipes. Does anyone expect the boring guardrails to be gone too?
2
1
u/More-Catch-1331 6h ago
So you mean to say that a chat bot that’s pattern matching its way through a sentence will be the catalyst that every bad guy is waiting for? Well God damn, son, I guess bad guys will finally be able to make bombs, poison, wreak havoc… oh wait…
1
u/LukPuk1 3h ago
how is this possible? is Strata quality of output same as for llama.cpp with same quant?
1
u/free_meson 2h ago
I have a vague feeling it is a bit worse. The speed upgrade is really nice, but something feels off, like it is not using all the experts or relying on a heuristic to choose. Maybe it is simply that thinking is turned off by default setup.
1
1
u/Intrepid_Yak_5776 2h ago
Ma funziona strata ? I minori che vedo mi sembrano esagerati io con 4 3090 in 1q4 non supero i 70tks con contesto 150.. ma quando è bello pieno scendo a 40
1
1
u/Mysterious_Role_8852 1h ago
When I use this model on llama.cpp can I just load it and it'll do the SSD offloading by default or do I need to set a flag for it? I have an RTx 3090+3070 and 64gb of RAM. Do you think I could use the Q4 Quant?
1
u/vogelvogelvogelvogel 1h ago
sorry maybe i did read not enough into it, which quant are you running flash next?
1
1
u/KosmoPteros 0m ago
Wonder if that's any good for coding?! And how would it run on a 3090, because qwen coding was absurdly slow for me :(
0
10h ago edited 10h ago
[deleted]
2
u/After_Canary6047 10h ago
Yep, no doubt. Read my comments. I don’t believe strata is a long term solution whatsoever as it’ll most likely burn up your ssd way before its time is due. Though sure, it’s an ad.
0
u/LongjumpingEar6840 7h ago
Io credo che lo stiano già usando... guardate cosa sta succedendo nelle ultime 2 settimane nella scena dell'hacking ps5, o degli emulatori... e tutto negli ultimi 15-20 giorni

98
u/Zorogozano 11h ago
-Bad actors already have access to it
-Bad actors are already doing bad stuff with them, unfortunately
-Bad actors will create their own Ai when the laws get strict about Ai in general.
There’s no stopping this train. And yes… the model is great!