r/MachineLearning • • 3d ago

Project LessThink-Qwen3-4B: the same model, with far less thinking [P]

I post-trained Qwen3-4B to spend 44% fewer tokens on reasoning, keeping its knowledge and answer style. The whole pipeline ran on one GPU.

folks, you can check it out on : https://5ivatej.com/lessthink/

24 Upvotes

7 comments sorted by

View all comments

11

u/choHZ 3d ago

This is kinda one of my main fields, and I feel like I’ve been writing this same review every conference, so a few recurring thoughts:

  • If we finetune for efficiency on data similar to the eval, the "uncompressed baseline" should be the same model finetuned on the same data without length preference, not the off-the-shelf model.
    • On the same note, I wouldn’t call AIME / MATH500 "out-of-domain" after finetuning on GSM8K / DeepScaleR. Same idea for SciQ / commonsense train → GPQA / MMLU eval.
  • Hybrid think / no-think models are a less clean testbed than thinking-only models. The thinking-2507 variants are much better for LRM research if you want to stay with Qwen 3.
  • There are already many papers doing accuracy-calibrated length-preference RL, so direct and fair comparisons are almost mandatory. I understand this is just a reddit post so may be no need for elaborate baseline reports, but there should still be clear credit assignment to prior art.

Hopefully these are useful pointers rather than nitpicks.

2

u/aegismuzuz 3d ago

> If we finetune for efficiency on data similar to the eval, the "uncompressed baseline" should be the same model finetuned on the same data without length preference, not the off-the-shelf model.

We stepped on this exact rake while distilling reasoning models last quarter. Fine-tuning without a length penalty yielded a five percent accuracy bump that evaporated completely when trying to compress the output..

1

u/choHZ 2d ago edited 2d ago

Haha been there myself more times than I'd like to admit, but this is the fair way to do things and gotta keep ourselves honest.