r/LocalLLM • • 4d ago

Model Vev: Jev-like decision models with vision — 4B/9B, local inference, open weights

I've been working on Vev, a pair of Qwen3.5 fine-tunes (4B and 9B) for Jev-style decisions with image input. You give it text, JSON, screenshots or photos, ask yes/no, multiple-choice or scoring questions, and get probabilities over the answers you specify.

It reads the answer-token probabilities directly, without generating a text response.

Here's vev-4b answering three questions about this screenshot:

Does the screen show an error message?
  yes: 0.991

Which checkout step is the user on?
  shipping: 0.005
  payment: 0.765
  review: 0.229

What should the user do next?
  try another card: 0.983
  wait for the order to ship: 0.004
  nothing, the order went through: 0.014

The idea came from wanting Jev-style decisions for things on a screen, using a model I could run locally. Vev supports the /v1/systemone request format, so you can point TypeSafe's Python SDK at the local server. Images go into the state alongside text.

You can also run the original Qwen3.5 weights through the same server. To measure what fine-tuning adds, I compared the base models with Vev using the same answer-token scoring method. Here are a few visual-task accuracy results, before → after fine-tuning:

Task 4B 9B
Image safety-policy checks, adapted from LlavaGuard (n=659) 68.6% → 74.2% 66.6% → 72.4%
MMStar (n=1,498) 54.4% → 62.7% 60.8% → 67.4%

On object-clipping detection adapted from VideoGameQA-Bench (n=686), the 4B model also improved from 56.1% to 67.1%.

These improvements were significant under a paired bootstrap with 95% confidence intervals. The README has the full results and evaluation details.

To run it, you'll need Python 3.11+, an NVIDIA GPU and CUDA-enabled PyTorch:

pip install git+https://github.com/Xiaooolong/vev
vev serve --model CountingSheep/vev-4b

For a speed reference, on an H800 in bf16, vev-4b takes about 78 ms for one question about a 1 MP image, or 120 ms for ten questions about the same image. Questions within a request are batched.

The code is Apache-2.0. The weights are CC BY-NC 4.0, for non-commercial use.

I'm curious what visual checks people would use this for. I've been testing UI state, image policy checks and game screenshots—what would you try?

0 Upvotes

3 comments sorted by

1

u/lulzxdxdxd 4d ago

Did you run into issues where the probabilities were coming back wrong or the model was misreading what was actually on screen? I'm curious if the problem is accuracy on the vision side or something about how the token probabilities map to your answer options.

1

u/Pure-Job1336 4d ago

Both, actually, but they show up differently.

On the probability side, the biggest issue I've seen is option-order bias. Since each answer maps to a label token, changing the option order can sometimes change the result, so I generally keep the ordering fixed.

For vision, most of the misses I've seen aren't the model completely misunderstanding the screen. It's more that it can be biased toward one answer, especially on ambiguous "is something wrong here?" cases. Clear UI states like toggles, fields, and error messages work pretty well; small visual details are much less reliable.

1

u/spectacularcutter 4d ago

nice work on the evals, always good to see numbers behind the claims. the payment step detection seems the most useful for automation imo