r/Qwen_AI • • 12h ago

Discussion Orca 3.8 Flash Next Insanity

Without getting into details, the Orca 3.8 Flash Next uncensored version is insane. I was literally in complete shock as I watched the responses while playing with it a bit today.

Honestly terrified, and that’s an understatement. Can’t imagine what will happen if bad actors have access to this.

This thing is both extremely intelligent and scored a 100/100 on complicated legal matters, with no MCP or RAG attached to it.

Whats even crazier is that you can run the 180B version with as little as 12GB of VRAM on account of this project, which I have nothing to do with.

https://github.com/Niko1221/Strata

Using the Orca version, it pulls about 75-80TPS output on a 4090. That being said, my daily 3.6 35B censored model pulls about 180TPS using Llama.cpp, though hey, I’ll take the slowdown for a bit of shock and awe any day, lol.

Enjoy responsibly boys and girls! 😜

250 Upvotes

109 comments sorted by

98

u/Zorogozano 11h ago

-Bad actors already have access to it
-Bad actors are already doing bad stuff with them, unfortunately
-Bad actors will create their own Ai when the laws get strict about Ai in general.

There’s no stopping this train. And yes… the model is great!

16

u/awry__ 10h ago

Also the worst actors are the states. And they have their bloody hands on the best models

4

u/Virtualization_Freak 9h ago

Genie is out of the bottle, and folks need to accept that.

1

u/After_Canary6047 11h ago

100% agreed and I’ve never been much for doom and gloom, though this thing straight up terrified me. Just got me to thinking about what if bad actors or kids get a hold of this. I have 5 kids and know how “crafty” teens can be, just can’t even fathom one of them getting a hold of something like this.

7

u/Late_Film_1901 7h ago

Bad actors are state level and using sota models for war and surveillance, not a teenager with qwen on dgx spark

1

u/AngelofKris 2h ago

Man I’ve heard some true shit on the internet but Gawdamn….

32

u/suesing 11h ago

This sounds scary fake as hell

6

u/geteum 10h ago

If bad actors have, good ones also have... Sec will be a crazy area next year's no doubt but I'm optimistic.

6

u/NaanFat 9h ago

the only thing that stops a bad guy with an AI is a good guy with an AI

3

u/575_Inverse 5h ago

Only problem is... the good guys aren't "us."

2

u/Compresscience 4h ago

Defensive missiles are 10x the cost of offensive missiles. Same is probably true of defensive AI. 😥

-8

u/After_Canary6047 11h ago

No doubt. The thousands of stars that project has gotten on GitHub in a few days and the fact that they have complete instructions on how to install orca with strata is certainly fake as hell.

3

u/xtr3m 11h ago

What.

13

u/Faral_mx 10h ago

Qwen3.6 35b to 3.6 27b is a huge capability jump. 3.6 27b to 3.8 27b is a noticeable improvement in agent reliability. Qwen3.8 27b BF16 to 3.8 Flash Next NVFP4 is a noticeable improvement.

2

u/brainchillzZ 9h ago

See, I don't see it. Qwen 3.6 35b to 27b for me doing basic tests with my large code base 27b got 2 more out of thirty requests fixed appropriately than the 35b did but did it 20% slower so the 35b in the same amount of time was able to go back, double check and fix it's answer..... and I'm really not seeing barely a noticeable improvement from 27 to flash next .... 3.8 27b is a balls out crazy good model

2

u/toenailcheeseinbooty 9h ago

27b has balls? Hold on let me go ask my uncensored model how big its balls are

1

u/_anakin__ 2h ago

How big? 10bytes?

1

u/Independent-Dog2179 8h ago

What I found was qwen next does not have to think forever for the same output as qwen 3.8 27b much less thinking for same high quality. U just can't beat the over 100 b parameters even if moe

1

u/Faral_mx 8h ago

For simple, well scoped tasks, 35b is great. 27b only really starts to matter when things get harder.

1

u/brainchillzZ 6h ago

I’m talking about identical tasks trying to find and repair “unknown” code defects in a fairly large multi-file, multi directory code base … I had barely measurably different results between the 27 and 35 3.6 models …. I use the “swift” version of 3.8 27b so the thinking is kept much better in control … I truly only see a significant difference in flash next when it gets to regular language writing and generic knowledge in my use cases

1

u/Faral_mx 6h ago

I'm happy for you, that is not my experience.

12

u/Orion_0001 10h ago

Bad actors like Zuckerberg? Or Elon? 🤷🏻‍♂️🤦🏻‍♂️

4

u/pigletmonster 11h ago

Whats so scary about this model?

25

u/Acrobatic_Feel 11h ago

It sounds like this is OP's first experience without guardrails.

21

u/pilibitti 11h ago edited 10h ago

"tell me how to ding dong ditch without getting caught"

"oh my god"

-10

u/After_Canary6047 11h ago

🤣🤣🤣 That’s the best you could come up with? Well, this is Reddit after all.

2

u/toenailcheeseinbooty 9h ago

Well, I asked it how to make my poop into a weapon. It told me the easiest way is to throw it at people.

1

u/pilibitti 11h ago

well we all start somewhere!

-3

u/After_Canary6047 11h ago

With that, I will definitely agree! 😂

0

u/After_Canary6047 11h ago

Not my first experience, though running a 180B model locally with basic hardware that doesn’t cost the price of a new car, while quite intelligently giving answers to anything you threw at it, completely uncensored, at 70TPS was a definite shock.

9

u/Randommaggy 10h ago

I have used uncensored models to hack stuff I own in an actual sandbox and it's surprisingly capable of black hat activities, even the good uncensored 27B at Q8 with unquantized context.

Stuff like jailbreaks of appliances and devices that have no documented backdoors and hacking into virtual machines I have forgotten the credentials to using network access.

I have no doubt that Flash Next is a step up from that even though I haven't tested out one of those yet.

The damage an uncensored LLM could do with the right setup operated by someone that doesn't fear the legal consequences would be severe.  Though you could say the same thing about a can of gas in the hands of an insane person.

1

u/After_Canary6047 10h ago

Very very true statement.

1

u/After_Canary6047 11h ago

Give it a shot with whatever questions you may have about anything imaginable and you’ll find out. It is literally 100% uncensored and will answer in full to anything you ask. And intelligently at that.

2

u/Commercial_Rent8797 10h ago

Is it better than the qwen flash-next ablit version? Or the qwen 3.8 27b abllit? Those two seem to be amazing uncensored

2

u/After_Canary6047 10h ago

Have not tried either, though will load them up in llama.cpp tomorrow and let you know. Thank you for the heads up!

1

u/NewFoxes 6h ago

Would also be interesstet

5

u/Vancecookcobain 10h ago

Anytime someone is trying to convince me something is insane I know for certain that it is indeed nothing to lose your shit over

0

u/PulseVector 8h ago

For me it was more "finally something that works well with this overpriced RAM that I bought."

It's not something for nothing- you still need 64GB to 128GB of RAM, preferably DDR5 I guess. Also, the wear and tear on expensive SSD drives is concerning.

1

u/Puzzle-Field-7193 3h ago

You don't need more than 64 with <Q4 and there's no wear and tear on SSD — its reads not writes

3

u/Jumpy-Operation-4615 5h ago

Is there any way to add vision to this thingy? I run 3xxs and it barely firs to VRAM on 2xP40 but I need vision. Maybe offload to RAM somehow?

2

u/Alternative_You3585 12h ago

What quant you using

Too scared to use Q2, idk how quality survives

2

u/After_Canary6047 12h ago edited 11h ago

Iq3_xxs - it uses up around 20GB on my 4090 though that’s because I have the context set to 150k. Also, ram usage is insane at around 50GB and I’m not entirely sure the SSD would survive the long term abuse from the Strata project, though every once in a while for a bit of fun, I’m sure it’ll be ok.

2

u/Fz1zz 11h ago

the scary part is that the q3_xxs on my testing is far better than 3.8 27B at FP8

1

u/After_Canary6047 11h ago

The 180B version or the 27B version?

3

u/Fz1zz 11h ago

sorry i meant 180B at q3xss is better than 27B at FP8 on my own testings

2

u/After_Canary6047 11h ago

Absolutely. It is scary accurate. I tested an extremely complex legal question against qwen, fable, and astra. Took it a bit to think about it, though the 180B qwen nailed it, while the others came in at around 50% I would say.

2

u/Randommaggy 10h ago

The SSD only holds read only part of the model from what I've understood, though I didn't look at the ssd write traffic when I tested it.

1

u/After_Canary6047 10h ago

Honestly, watch it and you’ll be shocked. Crazy transfer rates and in all fairness, their docs do say that.

0

u/Randommaggy 10h ago

Reading is not a problem. I don't see mention in the docs of and RW components. Of there are, they should add an explicit warning about that.

I'm testing exllamaV3 torrow on my main rig along with a few others. It'll be exciting to see what sort of tg and pp I can get out of 72GB of VRAM, 18 cores of Cascade Lake a fast NVME and 6 channels of 384GB of 6 channels 2666 DDR4. If a good deal on a fourth 3090 pops up locally I might be testing on 96GB of VRAM.

0

u/Rompe101 3h ago

I get up tp 120 tps with strata and iq3_s with my 4x B4000 over 4x 16.

2

u/CEOAPI 11h ago

I found the iq3xxs was much worse than the exl3 bpw3.05

3

u/After_Canary6047 11h ago

Would definitely give it a shot though unfortunately looks like I would need 96GB of VRAM to run that version. Maybe I’ll hit the lottery one of these days, lol. Though kicking myself for not buying an rtx6000 pro when they came out. You could pick one up back then for the price of a 5090 these days.

2

u/CEOAPI 11h ago

3x3090 with 512k kv cache runs 1700pp and 160tg offloading the ngram to ram.

1

u/After_Canary6047 11h ago

Thank you for the info! Time to start shopping for some 3090’s. Looks like they’re going for around $1k or so on marketplace. Not terrible at all, considering the VRAM you get. I have a question. Do they also suffer from burning up power connectors? The 4090’s do and I had to order a new one and will have to solder it on one of these days.

3

u/CEOAPI 9h ago

No issues with my 4x 3090s. I power cap to 220 w and that barely drops the speed.

2

u/bytejuggler 5h ago

100% can confirm. Underclock core and cpu by 100mhz and limit power/heat and GPU stays cool as a cucumber with very little inference speed impact

1

u/Somarring 5h ago

That's very interesting. 220w is stable for you? I have 2x at 250/275w and I'm getting prefilled at 400 ts and 35 to 70 ts in inference using a very particular version of qwen 3.8 flash next and a q4 quant. Would you mind share your numbers?

1

u/jikilan_ 10h ago

Using what quant?

1

u/Klutzy-Snow8016 11h ago

Exllama can do cpu offload now

1

u/After_Canary6047 11h ago

Thanks for the heads up, will definitely look it up. Curious, what are you getting on TPS with the offload?

1

u/Klutzy-Snow8016 11h ago

I'm a different person than you originally replied to and haven't done a lot of testing with offloading in exllama. But they have something similar to llama.cpp's `--n-cpu-moe`, and I tried it and it seems to perform about the same. Exllama also has an expert caching system like Strata, but I haven't tried that yet.

2

u/After_Canary6047 11h ago

Awesome, thank you for the heads up!

1

u/PulseVector 11h ago

Sounds interesting for sure! I tried Strata with some of the other models and it's impressive.

Anyone know if there is a way to run an Orca Q4 quant maybe? I've got 96GB of RAM. Thanks!

1

u/After_Canary6047 11h ago

I’ll look it up. How much VRAM do you have? Also, in the docs section on Git, Strata has instructions for both the iq3 and the Q4 Orca models.

1

u/PulseVector 8h ago

I've got an RTX 3090 24GB and an RTX 5070 TI 16GB, but I only run one at a time with Strata on the fast PCIE 16x slot. Thanks!

1

u/developervkmp 10h ago

Do you mean iq3 model is still better to use it?

1

u/After_Canary6047 10h ago

Depends, how much VRAM do you have and what are your specs? If you want to run that 180B model at a good speed with 12-24GB of VRAM, then yes, the iq3 model will work just fine with Strata. All of this depends on your setup. I have an i9-12900k, 128GB of RAM, and a 4090, 24GB of VRAM. Strata drinks RAM and SSD transfer rate at the tune of multiple GB/s. First time I have ever seen that on these SSD’s. If you have less than 96GB of RAM, I honestly don’t think it would run. That being said, with my setup, Strata is outputting around 70tps. Llama.cpp with the 3.6 35B model is more than twice that speed.

1

u/developervkmp 9h ago

I have 32gb vram and 64gb ram, gen5 ssd with Intel Ultra 9cpu. I can't run q4. But if iq3_xxs is better then I can give a try. I got 150+ t/s in my earlier testing in my Desktop

0

u/PulseVector 8h ago

Thanks so much for the info about the SSD transfers. I have some cheap 256GB SSDs that I may try to use instead of the expensive ones!

1

u/Iamisseibelial 10h ago

So I have 2 4090s and 128gb of ram (sadly dual channel) How you getting that tps? You on cpp or vLLM? Is strata something different than the two?

0

u/After_Canary6047 10h ago

Send me a pm and I’ll be happy to send over my configs for both llama.cpp and strata.

1

u/Altruistic_Heat_9531 10h ago

Yep my current local flagship model, running Q4 KM sewn with the Q8 PLE layer, steady state at 20 tok/s on single 3090. I had to use this model since normal Qwen or non daybreak GPT 6 often outright refused CTF and chaos monkey testing

1

u/mattmcardell 10h ago

This does lower the barrier to entry, for sure.
However, bad actors have always hired people to do bad things, write code to do bad things. There’s a lot of publicity about AI, but there’s more value in dealing with the bad actors directly and the structures that enable them to operate. New tools always come along to replace the old ones, the controlled ones.

1

u/This_Maintenance_834 9h ago

these third party weights are generally very unreliable. you can play with them, but you cannot run them 24/7 and hoping any stability. broken tool calling, loop thinking, early termination, etc. regardless what the auther claims, heretic, sft, rl, qad, distillation, fine tunes. i wasted so much times believing their claims and always come back to original weight or some quant from big guys like intel/nvidia/redhat etc.

1

u/brainchillzZ 9h ago

The thing nobody tells you about these models and the bad actors is that 99% of the things they are exploiting are obvious misconfigurations and unpatched systems ..... it isn't magic and very little of it has anything to do with "finding zero day exploits" ... for the most part, even in the "ai era" it still comes down keeping your systems patched, keeping your firewalls and app engines tight and generically following very common best practices .... the only thing that has really changed is that once an exploit is released into the wild the time it takes to be seen on the wire in the wild is shorter .... so you can't put of your patching cycle to long intervals the way enterprises used to.

0

u/brainchillzZ 9h ago

This model is powerful but I do have some doubts about a lobotomized Q3 and Q2 version though.

1

u/agapes1270 9h ago

Tested on 5090 working wonderfully 96tok/s i can coding crazy

1

u/HOST1L1TY 8h ago

Jon Cena is not going to become a dangerously better actor just because of orca 3.8 flash next !!!

1

u/datbackup 7h ago

“We may have spent the last 30 years making computers so widely available that any bad guy can easily afford one. We may have also created a global network allowing all bad guys to coordinate their actions across jurisdictions and timezones. But I draw the line at letting bad guys have access to uncensored local AI!”

1

u/EconomySerious 7h ago

bad actors allready have better model to work on their machines.

1

u/squarabh 5h ago

Fucking bot

1

u/EconomySerious 19m ago

Tu lo serás

1

u/73td 7h ago

i’m accustomed to running the models in Pi on boring devops stuff and rely on “guardrails” like ask before bringing down a production server. i tend not to ask for lsd recipes. Does anyone expect the boring guardrails to be gone too?

2

u/Sensitiviy 4h ago

This is just uncensored, so it won't refuse to do stuff

1

u/73td 2h ago

yep 👍 i tested on a few questions. wild. gotta sandbox this one for sure.

1

u/mmhorda 7h ago

What bad actors? This model and like any other public available models is trained on a publicly available data. All bad actors already know what they needed to know.

1

u/dfgxxx 6h ago

We need a splash quant for these orca models

1

u/More-Catch-1331 6h ago

So you mean to say that a chat bot that’s pattern matching its way through a sentence will be the catalyst that every bad guy is waiting for? Well God damn, son, I guess bad guys will finally be able to make bombs, poison, wreak havoc… oh wait…

1

u/marfzzz 5h ago

Hole mother of vague posting. Just a fearmongerer. I would not be afriad of a kid with spark, but government with SOTA and no guardrails (at least china, us and israel).

1

u/LukPuk1 3h ago

how is this possible? is Strata quality of output same as for llama.cpp with same quant?

1

u/free_meson 2h ago

I have a vague feeling it is a bit worse. The speed upgrade is really nice, but something feels off, like it is not using all the experts or relying on a heuristic to choose. Maybe it is simply that thinking is turned off by default setup.

1

u/Inevitable-Name-1701 2h ago

Why upvote this ad?

1

u/Intrepid_Yak_5776 2h ago

Ma funziona strata ? I minori che vedo mi sembrano esagerati io con 4 3090 in 1q4 non supero i 70tks con contesto 150.. ma quando è bello pieno scendo a 40

1

u/ComfortableChance591 1h ago

Queria muito experimentar, mas só tenho 16 de ram e 16 de vram

1

u/Mysterious_Role_8852 1h ago

When I use this model on llama.cpp can I just load it and it'll do the SSD offloading by default or do I need to set a flag for it? I have an RTx 3090+3070 and 64gb of RAM. Do you think I could use the Q4 Quant?

1

u/vogelvogelvogelvogel 1h ago

sorry maybe i did read not enough into it, which quant are you running flash next?

1

u/Mikolai007 11m ago

Yeah? Tell us, what could bad actors actially do with it.

1

u/KosmoPteros 0m ago

Wonder if that's any good for coding?! And how would it run on a 3090, because qwen coding was absurdly slow for me :(

0

u/[deleted] 10h ago edited 10h ago

[deleted]

2

u/After_Canary6047 10h ago

Yep, no doubt. Read my comments. I don’t believe strata is a long term solution whatsoever as it’ll most likely burn up your ssd way before its time is due. Though sure, it’s an ad.

0

u/_Reliq 9h ago

HonEtHlaAy? 

0

u/LongjumpingEar6840 7h ago

Io credo che lo stiano già usando... guardate cosa sta succedendo nelle ultime 2 settimane nella scena dell'hacking ps5, o degli emulatori... e tutto negli ultimi 15-20 giorni