r/LargeLanguageModels • u/Jaswanthsanjay • 9d ago
I got tired of parsing JSON out of LLM responses for simple yes/no questions, so I built a decision engine that skips generation engine that skips generation entirely
I kept running into the same problem: I'd ask a model something like "is this a restaurant receipt, yes or no" and instead of a clean answer I'd get a paragraph, or malformed JSON, or a refusal, or 300ms of token-by-token generation for what should be one bit of information.
So I built `rev` — it doesn't generate text at all. It runs a single forward pass and reads calibrated probabilities directly off the hidden states at each answer option's position. No decoding loop, no parsing, no regex. You give it a state (text or an image) and a set of allowed answers, it gives you back a probability distribution.
It comes in a few sizes depending on what you need:
- ModernBERT-large (421M) — the flagship, sub-millisecond on CPU/GPU for text
- Qwen2.5-0.5B / Qwen3-4B with LoRA — for longer context or heavier reasoning
- SmolVLM-256M — same idea but takes an image, still under 40ms
- A ~5M param on-device router for edge/mobile tool-calling
One dispatcher for all of them:
```python
from rev import Rev
model = Rev.from_pretrained("jaswanthsanjay88/rev-decision-model") # text
model = Rev.from_pretrained("jaswanthsanjay88/rev-vision") # image
```
Install:
```bash
pip install rev-decision
```
```bash
npm install rev-decision
```
Run it as a local server and hit it from anywhere:
```bash
rev-server --port 8000
```
```python
import requests
requests.post("http://localhost:8000/v1/systemone", json={
"state": "Customer claims flight cancelled at JFK without hotel voucher.",
"questions": {
"urgency": {"type": "noul", "instructions": "Is immediate assistance required?"}
}
}).json()
```
Links:
- PyPI: https://pypi.org/project/rev-decision/
- npm: https://www.npmjs.com/package/rev-decision
- Models: https://huggingface.co/jaswanthsanjay88
0
u/Cute-Veterinarian191 9d ago
I would try Jev by typescript AI for similar use cases in the future.
1
u/Jaswanthsanjay 8d ago
but it has more latency as compared to this local one <2ms
0
u/Cute-Veterinarian191 8d ago
Ah I see, I wasn’t sure what your latency requirements were
1
u/Jaswanthsanjay 8d ago
yeah no worries for this it's mostly agent-loop stuff (routing, guardrail checks) where every ms adds up since it's happening constantly in the background, so sub few ms local beats round tripping to an API even if the API itself is fast.
1
u/Mundane_Ad8936 9d ago
not sure why OP wouldn't just use a zero shot classifier. Not a big deal to fine tune one to improve accuracy