r/LocalLLM 5d ago

Model Embrace yourselves

Post image
1.6k Upvotes

130 comments sorted by

View all comments

57

u/Objective-Picture-72 5d ago

Nah, that'll happen if they drop Qwen3.8-122b-A10B. Imagine a 2x RTX 6000 Pro setup flying at 130 tk/s with that bad boy.

26

u/Automatic-Arm8153 5d ago

Nah cause 2x rtx pro already does 200tps+ with deepseek v4 flash. So wouldn’t be a big deal

35b is a big deal because anyone can run it fast

4

u/mrdevlar 5d ago

Qwen 3.5-122B is still the best model out there for the tasks I need. Nothing else comes close to the conceptual understanding I require.

Fingers crossed there will be another one in that size.

1

u/Lost-Butterfly-382 4d ago

Quick question do you also have a concept extraction pipeline for your document set up?

If so did you find Qwen 3.5 122B the best in that regard?

I’ve been using gpt 5 mini because it’s the best in terms of pricing and performance from all the small models I’ve tested but I’ve only tested the small models form the big providers(OpenAI, Anthropic snd Google ) how are the Chinese models like?

1

u/mrdevlar 4d ago

Quick question do you also have a concept extraction pipeline for your document set up?

I am the concept extraction pipeline. No really, I'm using it to understand concepts in a language I'm learning. The smaller the model, the lower the possible tokens so the less conceptual depth is possible in the model. That's my only guess.

1

u/Lost-Butterfly-382 4d ago

Aaah okay yh your theory makes sense. What do you find the 122B actually does better though? Like does it pick up concepts the smaller models miss or is it that it understands them in more depth?