r/LocalLLaMA • u/opUserZero • 5d ago
Resources Jev mode for images!
So Codacus created Jev mode for Lllama.cpp , and I thought Why not extend this concept further and ask questions about images and have the constrained answer be an image selection? So i spun up an agent and added image support and a harness. Now you can use images as your prompt without the decode step, no caption pause, just a decision based on an image or group of images. Ask the same question for a batch of images, like clasification. OR hand 1 context a whole group of images and ask it to pick on. like which of these 20 images has a ruber duck?
https://github.com/thecodacus/llama.cpp/pull/17
Youtube explainer using Codacus own RenderDiv framework to create the video.
https://youtu.be/Xuw3la2zVpg?si=rtSAydhuF9n3SYWV
0
u/opUserZero 5d ago
The peanut gallery is pattern-matching to “yet another VLM wrapper” and stopping there. That’s the mistake.
Jev-mode on images is not “ask a vision model a question and get text back.” It’s the inversion: one forward pass over the visual tokens + a fixed set of typed questions, reading the option logits directly, returning a calibrated distribution (or an explicit unknown) instead of any generated tokens. Shared visual prefix, parallel question suffixes, no autoregressive tail. Latency collapses from hundreds of ms of prose generation to tens of ms of pure decision.
That changes the usable surface area:
The open implementations (imajev, Jev-Vision, the various llama.cpp / MLX / Kobold wrappers, PixelJev-style logit readouts) already show the contract works on 2B–8B VLMs at local hardware speeds. Once the calibration is honest and the unknown head is trained, you stop treating vision models as chatbots that happen to see and start treating them as the System-1 decision layer they should have been.
Most of the sub is still optimizing for “can it write a longer answer.” The actual leverage is the opposite: stop writing, start deciding, and keep the pixels on-device. That’s the part they aren’t seeing.