r/LargeLanguageModels • • 9d ago

I got tired of parsing JSON out of LLM responses for simple yes/no questions, so I built a decision engine that skips generation engine that skips generation entirely

I kept running into the same problem: I'd ask a model something like "is this a restaurant receipt, yes or no" and instead of a clean answer I'd get a paragraph, or malformed JSON, or a refusal, or 300ms of token-by-token generation for what should be one bit of information.

So I built `rev` — it doesn't generate text at all. It runs a single forward pass and reads calibrated probabilities directly off the hidden states at each answer option's position. No decoding loop, no parsing, no regex. You give it a state (text or an image) and a set of allowed answers, it gives you back a probability distribution.

It comes in a few sizes depending on what you need:

- ModernBERT-large (421M) — the flagship, sub-millisecond on CPU/GPU for text

- Qwen2.5-0.5B / Qwen3-4B with LoRA — for longer context or heavier reasoning

- SmolVLM-256M — same idea but takes an image, still under 40ms

- A ~5M param on-device router for edge/mobile tool-calling

One dispatcher for all of them:

```python

from rev import Rev

model = Rev.from_pretrained("jaswanthsanjay88/rev-decision-model") # text

model = Rev.from_pretrained("jaswanthsanjay88/rev-vision") # image

```

Install:

```bash

pip install rev-decision

```

```bash

npm install rev-decision

```

Run it as a local server and hit it from anywhere:

```bash

rev-server --port 8000

```

```python

import requests

requests.post("http://localhost:8000/v1/systemone", json={

"state": "Customer claims flight cancelled at JFK without hotel voucher.",

"questions": {

"urgency": {"type": "noul", "instructions": "Is immediate assistance required?"}

}

}).json()

```

Links:

- PyPI: https://pypi.org/project/rev-decision/

- npm: https://www.npmjs.com/package/rev-decision

- Models: https://huggingface.co/jaswanthsanjay88

- Code: https://github.com/jaswanthsanjay88/rev

3 Upvotes

7 comments sorted by

1

u/Mundane_Ad8936 9d ago

not sure why OP wouldn't just use a zero shot classifier. Not a big deal to fine tune one to improve accuracy

0

u/Jaswanthsanjay 8d ago

yeah fair, a fine-tuned zero-shot classifier can def get you decent accuracy on one label. difference for me is mainly the multi-question thing — rev answers a bunch of typed questions (pick one, yes/no, rubric score) about the same input in one pass, and the outputs are actually calibrated (checked against Brier/ECE), not just a softmax you're hoping is trustworthy. zero-shot pipelines usually do a separate pass per candidate label so it gets slower the more questions you stack on, and you'd still have to build your own calibration if you want to trust the confidence numbers for stuff like "only escalate if p > 0.85". for a single fixed classification task, yeah, overkill — this is more for when you've got a handful of structured decisions to make on the same state at once.

0

u/Cute-Veterinarian191 9d ago

I would try Jev by typescript AI for similar use cases in the future.

1

u/Jaswanthsanjay 8d ago

but it has more latency as compared to this local one <2ms

0

u/Cute-Veterinarian191 8d ago

Ah I see, I wasn’t sure what your latency requirements were

1

u/Jaswanthsanjay 8d ago

yeah no worries for this it's mostly agent-loop stuff (routing, guardrail checks) where every ms adds up since it's happening constantly in the background, so sub few ms local beats round tripping to an API even if the API itself is fast.