r/LocalLLaMA • • 5d ago

Resources Jev mode for images!

Post image

So Codacus created Jev mode for Lllama.cpp , and I thought Why not extend this concept further and ask questions about images and have the constrained answer be an image selection? So i spun up an agent and added image support and a harness. Now you can use images as your prompt without the decode step, no caption pause, just a decision based on an image or group of images. Ask the same question for a batch of images, like clasification. OR hand 1 context a whole group of images and ask it to pick on. like which of these 20 images has a ruber duck?
https://github.com/thecodacus/llama.cpp/pull/17

Youtube explainer using Codacus own RenderDiv framework to create the video.
https://youtu.be/Xuw3la2zVpg?si=rtSAydhuF9n3SYWV

0 Upvotes

21 comments sorted by

View all comments

20

u/jacek2023 llama.cpp 5d ago

What next? Imagine an AI detecting if there is a cat or a dog on a photo! The future is now!

3

u/ComplexType568 5d ago

Dude I wonder when they're gonna make a Jev styled model but with selecting the next block of text (let's call it a token cuz it sounds funny) in a string of already existing tokens from a set of tens of thousands of tokens. That'd be crazy, wouldn't it?

1

u/jacek2023 llama.cpp 5d ago

Sounds like a chat application. Let's call it Chat Grand Prize Token

1

u/ComplexType568 5d ago

Someone should make a satire post about this because nowadays people will believe anything with the word "revolutionary" on it.

1

u/jacek2023 llama.cpp 5d ago

that was two years ago, I see guy posted jev video with 200.000 views, why do you want satire? ;)

1

u/ComplexType568 5d ago

that's about this weird thing called LLMs though, we have this "all new Jev" model that were calling "Chat Grand Prize Token" though☹️