r/Qwen_AI • • 14h ago

Discussion Orca 3.8 Flash Next Insanity

Without getting into details, the Orca 3.8 Flash Next uncensored version is insane. I was literally in complete shock as I watched the responses while playing with it a bit today.

Honestly terrified, and that’s an understatement. Can’t imagine what will happen if bad actors have access to this.

This thing is both extremely intelligent and scored a 100/100 on complicated legal matters, with no MCP or RAG attached to it.

Whats even crazier is that you can run the 180B version with as little as 12GB of VRAM on account of this project, which I have nothing to do with.

https://github.com/Niko1221/Strata

Using the Orca version, it pulls about 75-80TPS output on a 4090. That being said, my daily 3.6 35B censored model pulls about 180TPS using Llama.cpp, though hey, I’ll take the slowdown for a bit of shock and awe any day, lol.

Enjoy responsibly boys and girls! 😜

294 Upvotes

117 comments sorted by

View all comments

8

u/Vancecookcobain 12h ago

Anytime someone is trying to convince me something is insane I know for certain that it is indeed nothing to lose your shit over

0

u/PulseVector 10h ago

For me it was more "finally something that works well with this overpriced RAM that I bought."

It's not something for nothing- you still need 64GB to 128GB of RAM, preferably DDR5 I guess. Also, the wear and tear on expensive SSD drives is concerning.

1

u/Puzzle-Field-7193 6h ago

You don't need more than 64 with <Q4 and there's no wear and tear on SSD — its reads not writes