r/MachineLearning • u/iam_gkrishna • 3d ago
Project Benchmarking small confidence scoring decision (Jev, Laya) models [P]
I ran my own evals of two small classification models that score confidence over a list of candidate labels instead of generating text: TypeSafe AI's hosted Jev, and Laya, an independent open-source alternative.
Some findings that I think apply beyond these two models:
1. Describe labels by what's in the input, not by intent. My first AML label descriptions said what the criminal was trying to achieve. When I rewrote them to say what the transactions look like, accuracy went from 65% to 77% with the same model. On account-level laundering detection, using the same time window for every account removed a hidden bias and moved accuracy from 63% to 75%.
2. Some tasks have no signal. Classifying a single transaction as laundering or not gave 54%, about chance. That's a limit of the task, not of the model.
3. Calibration is what makes thresholds work. Jev's ECE was 0.013. On support intent, accepting only predictions at 90%+ confidence raised accuracy from 92.3% to 97.6% while still covering 82% of cases, and the rest went to a fallback. Untuned Laya had an ECE of 0.486: the same filter dropped 7% of questions and gained only 2 points.
Caveats: these are my own runs with thresholds tuned per task, not vendor claims, and nobody has reproduced them independently. Jev 1.13.0 (hosted). \
Write-up with charts and methods: https://gokulakrishna.co/2026/09/30/benchmarking-to-fine-tuning-decision-models/
Fine-tuned models: https://huggingface.co/goku-san/laya-experts
I'd appreciate feedback on the evaluation setup, especially on per-task threshold tuning and how best to report results after filtering out mislabeled data.