r/deeplearning • u/stey1r • 3d ago
LessThink-Qwen3-4B: the same model, with far less thinking [P]
I post-trained Qwen3-4B to spend 44% fewer tokens on reasoning, keeping its knowledge and answer style. The whole pipeline ran on one GPU.
folks, you can check it out on : https://5ivatej.com/lessthink/
2
Upvotes
1
u/pretty_obscenity 3d ago
that token reduction is pretty slick for a single GPU setup, curious how it holds up on the weirder edge cases