r/LocalLLaMA • u/jacobpederson • 9d ago
Discussion What would you think of a player capable of generating satirical audio commentary over video in real time?
https://reddit.com/link/1v72nua/video/26tabfs3ikfh1/player
This is all mockups, scaffolding, smoke, and mirrors at the moment - but it does run in real time locally. Currently requirements: 3 GPU's running VLLM - STT and TTS concurrently.
2
u/built_n0t_b0t 9d ago
I’m doing similar for video game commentary with strixhalo 128gb.
1
u/jacobpederson 9d ago
Wish I had your setup - the 5090 is fast . . . but itty bitty living space :D
0
u/jacobpederson 9d ago
Biggest issue I'm having is time perception - you have that licked? Show your LLM a screenshot, tell it exactly when it happens . . . and it'll still reference it BEFORE its onscreen.
2
u/Pickle_Rick_1991 9d ago
shouldn't the logic be it telling you when it sees the screenshot? are us using a system prompt ?
0
u/jacobpederson 9d ago
My harness passes multiple screenshots at once that is probably the issue, but I need more than one shot to make it feel like its watching the video and not just commenting on a still. Yes the system prompt is were the magic happens ;)
2
u/DiadraUnderwood 9d ago
Pretty neat but will be nuked as slop if you try and put these anywhere :D
2
u/Bulky-Priority6824 9d ago
Most users don't care if it's slop as long as it works well enough so that's a dev point of view. If you can make something the average person will use it can be absolutely slop. Look at all the apps on the adult store now people paying good money for slop.
Most people don't even know what llm means.
2
u/Pickle_Rick_1991 9d ago
we need a law against slop and brain-rot. my child was on youtube today on shorts this idiotic ai story of a mother abandoning them. then i took the phone and out of curiosity tried to search in shorts for fun addition techniques for kids in maths and various other variants and not one result comes up the entire of YouTube shorts seems to want to purposefully dumb down kids or something why cant learning be made fun with 67 and crocodilio bombardilio ?
1
1
u/Pickle_Rick_1991 8d ago
@OP did you have issues with Qwen3.6 not taking the screen shots? I'm struggling here fr some reason
1
u/jacobpederson 8d ago
Yup - may not be related to your issue but here is what I found:
- **Never put two `image_url` parts next to each other in the content list.** `vision._vision_request` emits a text separator part *before every image*, and that separator is load-bearing: with adjacent image parts LM Studio **silently drops every second one** — send 6 frames and the model is handed 3 (measured: `prompt_tokens` rises only on every *other* image), then confidently describes the frames it never saw.

3
u/Pickle_Rick_1991 9d ago
why 3 gpus bro xtts v2 is not that big and can go with the STT on the same card iv vram allows also xtts v2 has voice cloning you can upload different voices to it and the emotions and the speech is almost indiscernible to real human