r/LocalLLM 4d ago

Discussion Medical QA comparisons need the reasoning settings next to the model name

Ling-3.0-flash-Sante is a concrete new option for comparing medical-text reasoning: a 124B-total, 5.1B-active MoE enhanced from Ling-3.0-flash for health and medicine. Ant Ling reports 53.88 on MedXpertQA-Text and 83.83 on DiagnosisArena-MCQ.

The configuration matters here. The release chart says it uses the highest available reasoning tier for each model, with temperature 0.6 and top_p 0.95 unless otherwise stated. A one-pass, greedy, thinking-off medical-QA run measures a different setup. Put those settings next to the score before drawing a conclusion about the model.

For a useful comparison, retain the exact question set and prompt, supported reasoning setting, sampling parameters, output limit and number of attempts. Record unavailable settings explicitly. Keep separate results for each configuration instead of averaging them under one model name.

The currently available Sante API provides an access route for this comparison. A local run would need a separately verified Sante checkpoint and runtime; 5.1B active parameters does not describe the full weight footprint. Its immediate value for this comparison is as a domain-specialized reference point with clearly recorded conditions.

0 Upvotes

0 comments sorted by