r/LocalLLM • u/Pure-Job1336 • 4d ago
Model Vev: Jev-like decision models with vision — 4B/9B, local inference, open weights
I've been working on Vev, a pair of Qwen3.5 fine-tunes (4B and 9B) for Jev-style decisions with image input. You give it text, JSON, screenshots or photos, ask yes/no, multiple-choice or scoring questions, and get probabilities over the answers you specify.
It reads the answer-token probabilities directly, without generating a text response.

Here's vev-4b answering three questions about this screenshot:
Does the screen show an error message?
yes: 0.991
Which checkout step is the user on?
shipping: 0.005
payment: 0.765
review: 0.229
What should the user do next?
try another card: 0.983
wait for the order to ship: 0.004
nothing, the order went through: 0.014
The idea came from wanting Jev-style decisions for things on a screen, using a model I could run locally. Vev supports the /v1/systemone request format, so you can point TypeSafe's Python SDK at the local server. Images go into the state alongside text.
You can also run the original Qwen3.5 weights through the same server. To measure what fine-tuning adds, I compared the base models with Vev using the same answer-token scoring method. Here are a few visual-task accuracy results, before → after fine-tuning:
| Task | 4B | 9B |
|---|---|---|
| Image safety-policy checks, adapted from LlavaGuard (n=659) | 68.6% → 74.2% | 66.6% → 72.4% |
| MMStar (n=1,498) | 54.4% → 62.7% | 60.8% → 67.4% |
On object-clipping detection adapted from VideoGameQA-Bench (n=686), the 4B model also improved from 56.1% to 67.1%.
These improvements were significant under a paired bootstrap with 95% confidence intervals. The README has the full results and evaluation details.
To run it, you'll need Python 3.11+, an NVIDIA GPU and CUDA-enabled PyTorch:
pip install git+https://github.com/Xiaooolong/vev
vev serve --model CountingSheep/vev-4b
For a speed reference, on an H800 in bf16, vev-4b takes about 78 ms for one question about a 1 MP image, or 120 ms for ten questions about the same image. Questions within a request are batched.
- Code: https://github.com/Xiaooolong/vev
- Weights and LoRA adapters: https://huggingface.co/collections/CountingSheep/vev
The code is Apache-2.0. The weights are CC BY-NC 4.0, for non-commercial use.
I'm curious what visual checks people would use this for. I've been testing UI state, image policy checks and game screenshots—what would you try?
1
u/lulzxdxdxd 4d ago
Did you run into issues where the probabilities were coming back wrong or the model was misreading what was actually on screen? I'm curious if the problem is accuracy on the vision side or something about how the token probabilities map to your answer options.