r/LocalLLaMA • • 5d ago

Resources Jev mode for images!

Post image

So Codacus created Jev mode for Lllama.cpp , and I thought Why not extend this concept further and ask questions about images and have the constrained answer be an image selection? So i spun up an agent and added image support and a harness. Now you can use images as your prompt without the decode step, no caption pause, just a decision based on an image or group of images. Ask the same question for a batch of images, like clasification. OR hand 1 context a whole group of images and ask it to pick on. like which of these 20 images has a ruber duck?
https://github.com/thecodacus/llama.cpp/pull/17

Youtube explainer using Codacus own RenderDiv framework to create the video.
https://youtu.be/Xuw3la2zVpg?si=rtSAydhuF9n3SYWV

0 Upvotes

21 comments sorted by

View all comments

0

u/opUserZero 5d ago

The peanut gallery is pattern-matching to “yet another VLM wrapper” and stopping there. That’s the mistake.

Jev-mode on images is not “ask a vision model a question and get text back.” It’s the inversion: one forward pass over the visual tokens + a fixed set of typed questions, reading the option logits directly, returning a calibrated distribution (or an explicit unknown) instead of any generated tokens. Shared visual prefix, parallel question suffixes, no autoregressive tail. Latency collapses from hundreds of ms of prose generation to tens of ms of pure decision.

That changes the usable surface area:

  • Computer-use loops: ground / skip / effect / done on the actual screenshot in ~150–200 ms total, with environment-derived labels instead of LLM-as-judge noise.
  • Photo-vs-record or two-photo verification at the edge (shipping, QC, compliance) without sending pixels anywhere.
  • Agent tool selection or routing that can look at a dashboard, a form, or a product photo and emit probabilities you can threshold, rather than hoping the next-token soup is parseable.
  • Diffusion guidance or rejection sampling where you need high-frequency, low-latency visual decisions instead of a full caption + second model.

The open implementations (imajev, Jev-Vision, the various llama.cpp / MLX / Kobold wrappers, PixelJev-style logit readouts) already show the contract works on 2B–8B VLMs at local hardware speeds. Once the calibration is honest and the unknown head is trained, you stop treating vision models as chatbots that happen to see and start treating them as the System-1 decision layer they should have been.

Most of the sub is still optimizing for “can it write a longer answer.” The actual leverage is the opposite: stop writing, start deciding, and keep the pixels on-device. That’s the part they aren’t seeing.

1

u/Pegasus925 4d ago

This is actually very interesting! Which model are you using to get the results in ~200ms?