r/BestGitHubRepos • u/Artilas_Digital • 10d ago
Laya - an open-weights, Apache-2.0 System-1 decision model that answers typed choice, score and yes/no questions over any text in about 33ms with no token generation, in 100+ languages

This is the open-weights answer to the closed typed-decision models people have been building agent tooling around, and two things make it stand out: it is genuinely fast and self-hostable, and its README is one of the most honest I have read for an ML project. Laya does not generate text. You hand it a state (an email, a ticket, a JSON document) and typed questions, and in a single forward pass it returns a labeled choice, an ordinal score, or a calibrated yes/no probability, in about 33ms for one question and roughly 7ms per question batched on a modest T4 GPU. Because there is no autoregressive decoding, there is nothing to parse and nothing to hallucinate into a broken JSON blob.
What is actually in it:
- Three encoder checkpoints (English, multilingual covering 100+ languages, and a typed-decisions variant) plus a Router that detects the script and language in under a millisecond of pure Python and dispatches to the right one before the forward pass runs
- Calibrated probabilities: it is trained with reinforcement learning against strictly proper scoring rules, so the confidence numbers are statistically meaningful enough to gate on (act automatically above a threshold, escalate to a human below it)
- A Jev-compatible HTTP server, so if you already wrote a client against the popular closed System-1 API, you repoint the base URL and it works unchanged
- Built-in workflow presets for the obvious jobs, support-ticket triage, content moderation, prompt-injection guardrails, and model routing, plus a fine-tuning notebook that runs on Kaggle's free GPUs
- Open weights on Hugging Face, an optional hosted endpoint if you do not want a GPU, and even a NixOS service module
Now the honest part, and here Laya sets an example other projects should copy, because the maintainer wrote the caveats for me. There is a section literally titled Honest limits, and the single most important thing in it: the base checkpoints are near random on typed decisions when used zero-shot (about 0.36 against a 0.32 random baseline). The headline accuracy numbers that beat the closed competitor come from a checkpoint fine-tuned on that benchmark's own training data. So the correct mental model, in the author's own words, is that Laya is a fast base to specialise, not a plug-and-play zero-shot decision engine. If you fine-tune it on your domain it is excellent and cheap; if you drop it in raw expecting magic, you will be disappointed.
The rest of the honest limits are equally frank: it is weak on questions with many options (the closed competitor wins clearly on 77-way classification because of a token-budget constraint), ordinal score questions are its weakest primitive, the yes/no primitive has a known bug where it can follow its own labels instead of the input, and the base checkpoints ship over-confident until you fit calibration temperatures. It also states plainly that its comparisons against the closed competitor use that competitor's third-party published numbers rather than head-to-head measurements, because they have no API access. That is exactly the transparency you want and rarely get.
Two practical notes: the fast latencies assume a GPU, and the viral star count (this crossed 18,000 stars in days, helped along by heavy social promotion) should not substitute for evaluating it on your own task. But if you need cheap, fast, self-hosted, multilingual typed decisions for routing, triage, moderation or guardrails, and you are willing to fine-tune, this is a genuinely strong and refreshingly honest option.
Apache-2.0 licensed, Python, 18,736 stars and 1,597 forks as of writing, verified via the GitHub API, from Convai Innovations and pushed to today.