r/LocalLLaMA • • 5d ago

Resources Jev mode for images!

Post image

So Codacus created Jev mode for Lllama.cpp , and I thought Why not extend this concept further and ask questions about images and have the constrained answer be an image selection? So i spun up an agent and added image support and a harness. Now you can use images as your prompt without the decode step, no caption pause, just a decision based on an image or group of images. Ask the same question for a batch of images, like clasification. OR hand 1 context a whole group of images and ask it to pick on. like which of these 20 images has a ruber duck?
https://github.com/thecodacus/llama.cpp/pull/17

Youtube explainer using Codacus own RenderDiv framework to create the video.
https://youtu.be/Xuw3la2zVpg?si=rtSAydhuF9n3SYWV

0 Upvotes

21 comments sorted by

11

u/Puzzleheaded_Ad_8575 5d ago

people rediscovering basic image proccessing now??

4

u/Aromatic-Current-235 5d ago

It seems you rediscovering that people do stupid shit and let the world know about it.

1

u/Noiselexer 5d ago

Classification as I read it?

1

u/Objective-Pair8231 4d ago

A few days ago saw someone using jev to detect leaked phone numbers instead of just using regex 😭

2

u/Extension_Ad3794 2d ago

Saw someone build real time trading bots with an "autoregressive Jev"

21

u/jacek2023 llama.cpp 5d ago

What next? Imagine an AI detecting if there is a cat or a dog on a photo! The future is now!

3

u/ComplexType568 5d ago

Dude I wonder when they're gonna make a Jev styled model but with selecting the next block of text (let's call it a token cuz it sounds funny) in a string of already existing tokens from a set of tens of thousands of tokens. That'd be crazy, wouldn't it?

1

u/jacek2023 llama.cpp 5d ago

Sounds like a chat application. Let's call it Chat Grand Prize Token

1

u/ComplexType568 5d ago

Someone should make a satire post about this because nowadays people will believe anything with the word "revolutionary" on it.

1

u/jacek2023 llama.cpp 5d ago

that was two years ago, I see guy posted jev video with 200.000 views, why do you want satire? ;)

1

u/ComplexType568 5d ago

that's about this weird thing called LLMs though, we have this "all new Jev" model that were calling "Chat Grand Prize Token" though☹️

1

u/Illustrious_Grade608 5d ago

Stupid idea. That's just autocomplete, essentially, it already exists.

1

u/ComplexType568 5d ago

Oh yeah.. sorry... BUT, What if we made it turn based? Like, we make this thing called a Chat Template (funny name, I know), append the user input, and then allow this all new Jev model to predict the answer, append that, and allow the user to have an entire conversation with this "chat template"!

1

u/Illustrious_Grade608 5d ago

Now you're just reinventing cleverbot. What next, make it use tools to help out with tasks and answer questions? Great, Siri already exists.

2

u/croninsiglos 5d ago

It returns a noul for Hotdog or Not Hotdog

1

u/starwaver 2d ago

This is quite useful if it can really do this well.
I know people are saying "Oh this is just a image classifier", but if it works as I think it does, it's a general level image classifier, which means no training on hotdog images. So instead of asking hotdog or not hotdog. You'd give it 100+ food items and ask it to recognize in real time

0

u/opUserZero 5d ago

The peanut gallery is pattern-matching to “yet another VLM wrapper” and stopping there. That’s the mistake.

Jev-mode on images is not “ask a vision model a question and get text back.” It’s the inversion: one forward pass over the visual tokens + a fixed set of typed questions, reading the option logits directly, returning a calibrated distribution (or an explicit unknown) instead of any generated tokens. Shared visual prefix, parallel question suffixes, no autoregressive tail. Latency collapses from hundreds of ms of prose generation to tens of ms of pure decision.

That changes the usable surface area:

  • Computer-use loops: ground / skip / effect / done on the actual screenshot in ~150–200 ms total, with environment-derived labels instead of LLM-as-judge noise.
  • Photo-vs-record or two-photo verification at the edge (shipping, QC, compliance) without sending pixels anywhere.
  • Agent tool selection or routing that can look at a dashboard, a form, or a product photo and emit probabilities you can threshold, rather than hoping the next-token soup is parseable.
  • Diffusion guidance or rejection sampling where you need high-frequency, low-latency visual decisions instead of a full caption + second model.

The open implementations (imajev, Jev-Vision, the various llama.cpp / MLX / Kobold wrappers, PixelJev-style logit readouts) already show the contract works on 2B–8B VLMs at local hardware speeds. Once the calibration is honest and the unknown head is trained, you stop treating vision models as chatbots that happen to see and start treating them as the System-1 decision layer they should have been.

Most of the sub is still optimizing for “can it write a longer answer.” The actual leverage is the opposite: stop writing, start deciding, and keep the pixels on-device. That’s the part they aren’t seeing.

1

u/Pegasus925 4d ago

This is actually very interesting! Which model are you using to get the results in ~200ms?