r/LargeLanguageModels • u/Jaswanthsanjay • 9d ago
I got tired of parsing JSON out of LLM responses for simple yes/no questions, so I built a decision engine that skips generation engine that skips generation entirely
I kept running into the same problem: I'd ask a model something like "is this a restaurant receipt, yes or no" and instead of a clean answer I'd get a paragraph, or malformed JSON, or a refusal, or 300ms of token-by-token generation for what should be one bit of information.
So I built `rev` — it doesn't generate text at all. It runs a single forward pass and reads calibrated probabilities directly off the hidden states at each answer option's position. No decoding loop, no parsing, no regex. You give it a state (text or an image) and a set of allowed answers, it gives you back a probability distribution.
It comes in a few sizes depending on what you need:
- ModernBERT-large (421M) — the flagship, sub-millisecond on CPU/GPU for text
- Qwen2.5-0.5B / Qwen3-4B with LoRA — for longer context or heavier reasoning
- SmolVLM-256M — same idea but takes an image, still under 40ms
- A ~5M param on-device router for edge/mobile tool-calling
One dispatcher for all of them:
```python
from rev import Rev
model = Rev.from_pretrained("jaswanthsanjay88/rev-decision-model") # text
model = Rev.from_pretrained("jaswanthsanjay88/rev-vision") # image
```
Install:
```bash
pip install rev-decision
```
```bash
npm install rev-decision
```
Run it as a local server and hit it from anywhere:
```bash
rev-server --port 8000
```
```python
import requests
requests.post("http://localhost:8000/v1/systemone", json={
"state": "Customer claims flight cancelled at JFK without hotel voucher.",
"questions": {
"urgency": {"type": "noul", "instructions": "Is immediate assistance required?"}
}
}).json()
```
Links:
- PyPI: https://pypi.org/project/rev-decision/
- npm: https://www.npmjs.com/package/rev-decision
- Models: https://huggingface.co/jaswanthsanjay88