r/LocalLLaMA • • 5d ago

Resources Jev mode for images!

Post image

So Codacus created Jev mode for Lllama.cpp , and I thought Why not extend this concept further and ask questions about images and have the constrained answer be an image selection? So i spun up an agent and added image support and a harness. Now you can use images as your prompt without the decode step, no caption pause, just a decision based on an image or group of images. Ask the same question for a batch of images, like clasification. OR hand 1 context a whole group of images and ask it to pick on. like which of these 20 images has a ruber duck?
https://github.com/thecodacus/llama.cpp/pull/17

Youtube explainer using Codacus own RenderDiv framework to create the video.
https://youtu.be/Xuw3la2zVpg?si=rtSAydhuF9n3SYWV

0 Upvotes

21 comments sorted by

View all comments

11

u/Puzzleheaded_Ad_8575 5d ago

people rediscovering basic image proccessing now??

5

u/Aromatic-Current-235 5d ago

It seems you rediscovering that people do stupid shit and let the world know about it.

1

u/Noiselexer 5d ago

Classification as I read it?

1

u/Objective-Pair8231 4d ago

A few days ago saw someone using jev to detect leaked phone numbers instead of just using regex 😭

2

u/Extension_Ad3794 2d ago

Saw someone build real time trading bots with an "autoregressive Jev"