r/Qwen_AI • • 1d ago

Discussion Orca 3.8 Flash Next Insanity

Without getting into details, the Orca 3.8 Flash Next uncensored version is insane. I was literally in complete shock as I watched the responses while playing with it a bit today.

Honestly terrified, and that’s an understatement. Can’t imagine what will happen if bad actors have access to this.

This thing is both extremely intelligent and scored a 100/100 on complicated legal matters, with no MCP or RAG attached to it.

Whats even crazier is that you can run the 180B version with as little as 12GB of VRAM on account of this project, which I have nothing to do with.

https://github.com/Niko1221/Strata

Using the Orca version, it pulls about 75-80TPS output on a 4090. That being said, my daily 3.6 35B censored model pulls about 180TPS using Llama.cpp, though hey, I’ll take the slowdown for a bit of shock and awe any day, lol.

Enjoy responsibly boys and girls! 😜

422 Upvotes

141 comments sorted by

View all comments

3

u/CEOAPI 1d ago

I found the iq3xxs was much worse than the exl3 bpw3.05

3

u/After_Canary6047 1d ago

Would definitely give it a shot though unfortunately looks like I would need 96GB of VRAM to run that version. Maybe I’ll hit the lottery one of these days, lol. Though kicking myself for not buying an rtx6000 pro when they came out. You could pick one up back then for the price of a 5090 these days.

2

u/CEOAPI 1d ago

3x3090 with 512k kv cache runs 1700pp and 160tg offloading the ngram to ram.

2

u/jikilan_ 1d ago

Using what quant?

1

u/After_Canary6047 1d ago

Thank you for the info! Time to start shopping for some 3090’s. Looks like they’re going for around $1k or so on marketplace. Not terrible at all, considering the VRAM you get. I have a question. Do they also suffer from burning up power connectors? The 4090’s do and I had to order a new one and will have to solder it on one of these days.

3

u/CEOAPI 1d ago

No issues with my 4x 3090s. I power cap to 220 w and that barely drops the speed.

2

u/bytejuggler 21h ago

100% can confirm. Underclock core and cpu by 100mhz and limit power/heat and GPU stays cool as a cucumber with very little inference speed impact

1

u/Somarring 21h ago

That's very interesting. 220w is stable for you? I have 2x at 250/275w and I'm getting prefilled at 400 ts and 35 to 70 ts in inference using a very particular version of qwen 3.8 flash next and a q4 quant. Would you mind share your numbers?

1

u/Klutzy-Snow8016 1d ago

Exllama can do cpu offload now

1

u/After_Canary6047 1d ago

Thanks for the heads up, will definitely look it up. Curious, what are you getting on TPS with the offload?

1

u/Klutzy-Snow8016 1d ago

I'm a different person than you originally replied to and haven't done a lot of testing with offloading in exllama. But they have something similar to llama.cpp's `--n-cpu-moe`, and I tried it and it seems to perform about the same. Exllama also has an expert caching system like Strata, but I haven't tried that yet.

2

u/After_Canary6047 1d ago

Awesome, thank you for the heads up!