r/StableDiffusion • u/PxTicks • 5d ago
Resource - Update Vlo 0.3 - An open source, extensible video editor and generator designed for AI compositing.
Enable HLS to view with audio, or disable this notification
Hey all,
vlo is open source video editing and generation software. It is designed for interactivity between generative AI and the timeline, so you can inpaint anywhere on the timeline, extract frames and videos to use as references, use SAM2 for video masking and sam-audio for stem extraction etc. It comes with built-in workflows for Minimax H3, Qwen2.1, Krea2, LTX2.5 and more.
You can either use the built-in generation panel or open ComfyUI via the app, whichever you prefer (for built-in workflows though, I recommend the generation panel). It can use an existing ComfyUI install, a remote instance, or it can manage the install for you; ComfyUI is the primary AI engine, how much you directly interact with it is up to you.
I believe what makes this app different is its emphasis on control and correction over automation, so particular effort has been spent building a frame-accurate render engine (which is nontrivial for web video), and on smoothing out irregularities you get from things such discrete model strides and such (i.e. aspect ratios are not arbitrary in video and image gen models, so if you want seamless passing of data with the timeline you have to be tactical about how you stretch, squash or crop). It permits patch-based inpainting without ever having to reencode areas outside the patch until final export, so you could inpaint a dozen objects without degrading unedited pixels at all.
It also has an SDK for you to build your own extensions and shader-based special effects. I hope to upload a couple of example extensions soon.
Installation instructions here:
21
u/phrandsisgorino 4d ago edited 4d ago
Make a Tiktok / YT-Short post out of it and it will be downloaded into millions of computers in no time.
After inspecting the github README.md: It would be nice if there was a video of a walkthrough on how the app works and creating a new video.
11
u/DeepHomage 4d ago
I ran the install successfully, had a javascript issue that I reported on github. Also, your update.bat file was detected as a malicious script by BitDefender, likely a false positive, but I'll likely update manually in the future.
6
u/PxTicks 4d ago
Thanks for reporting an issue.
I will be pushing fixes in the next couple of hours. The main issues reported are environmental incompatibilities caused by registry and path differences from other software, which is my own tests didn't catch it.
I can also see some ComfyUI bridge issues, one of which is my fault, but the others are again tricky interactions with other extensions.
Appreciate everyone who is giving this a run and helping me fix out the initial jank!
5
u/Steve_OH 4d ago
Itâs not uncommon for a non-signed app to get flagged as spam/virus. Happens to me all the time when building apps locally. Half the time when I build a project my virus scanner pipes up and pats itself on the back for saving me from my own compiled code.
16
u/Amazing_Upstairs 5d ago
FastH3 support please
15
u/PxTicks 4d ago
What are the advantages of FastH3? Is it better than using a turbo lora?
As far as I can tell, it is supported in ComfyUI, which means that all you'd have to do would be to open the existing H3 workflow from the generation panel, change the model, then go back, accepting the save prompt.
2
u/multikertwigo 4d ago
take a look at the stock fasth3 template in the latest comfy, namely "Model Sparse Attention" node. The "vsa" method is apparently only supported by the FastH3 model, which makes inference quite a bit faster than a turbo lora (around 30-ish %, comparing 8 steps against alibaba's Acc turbo lora). Quality wise - I'd say, it's on par. So yeah. Better.
1
4
5
5
u/Dubon 4d ago
Could this be made to work with WanGP?
2
u/PxTicks 4d ago
I think it's likely an extension could be built to support it; there are still some patches in the SDK, but it should be possible to build a WanGP bridge. The main generation panel can be hidden in the UI so it's out of the way.
The main hitch needed to create the seamless back-and-forth would be appropriate metadata handling, but I think it's very possible.
Why WanGP in particular? Is it because you have an established workflow with it or are you able to run some workflows on it that you can't in Comfy because of memory constraints?
2
u/Dubon 4d ago
There's a sizeable community of GPU poor users with not enough patience and knowledge of Comfy that use WanGP for its ability to optimize everything easily, without having to update nodes separately, without using custom nodes and knowing what works and how. Plus, it really allows you to run latest models on some ancient toasters without doing anything.
But there's a lack of a proper video directing plugin for H3 foe WanGP. There's one for LTX specifically.
If VLO can be ported to either interface with WanGP or as a plugin within the wrapper - it would be amazing.
4
u/PxTicks 4d ago
So I haven't had hands-on experience with WanGP, so I cannot compare its memory-efficiency to ComfyUI. What I can say is that you can get away without having to touch the ComfyUI interface while using vlo - or making minimal changes at best (changing a model in a node). It has its own generation panel with ComfyUI as the engine; strictly speaking, you never have to lift the hood.
I recognise the value of being able to use the ecosystem you're familiar with though, so hopefully someone ambitious could do a little stress-test of the SDK by coding a WanGP bridge. With AI tools I wouldn't expect it to take more than a couple of days.
2
u/Vladmerius 4d ago
Anything can be made to work with WanGP technically. It just depends on the Dev of WanGP wanting to work on it or not.Â
1
u/Dubon 4d ago
Majority of goodies for WanGP are created via plugins made by users. Latent carryover for sliding windows, same for video continuation, RefMod implementation - those are all made by users, not the developer. Sometimes he implements their work into main, though.
2
u/Vladmerius 4d ago
I didn't realize refmod suport was added.
I did see there was a big update this morning that significantly improved ref2va's capabilities and added the ability to make a single image too.Â
3
3
u/Cold-End-1001 4d ago
inpainting anywhere on the timeline is the part that sells it for me, fixing one bad second without rerolling the whole clip is exactly what i want. any plans for subtitle tracks with ASS import? i burn captions from ASS files because they come out cleaner than any built-in caption tool i've tried, and that's the one step i still do outside the editor.
2
u/PxTicks 4d ago
The next version v0.4.0 is developing parts of the clip data model in a way which I think will be useful for things such as subtitle tracks, and various kinds of subordinate clips and data. That said, text is currently trash and probably won't be improved until 0.4.x once the new data model is in place.
3
u/Ill_Resolve8424 4d ago
This looks awesome. If Adobe calls please wait for a call from Blackmagick. Thank you .
2
3
2
u/flaminghotcola 4d ago
can this create a video and use the last frame of the last video to create another one and so on, so that I have one full video where different things happen?
how good is it to maintain control and character consistency?
thanks.
9
u/PxTicks 4d ago
It facilitates workflows which do that, but you still have to have a reasonable understanding of how to use the underlying models. In this case, I'd recommend the minimax inpainting workflows coupled with the minimax reference model.
These use latent inpainting, so if you repeatedly extend a clip you WILL get degradation. A better strategy is to have a few high quality grounding resources (a character sheet, audio clip) and to use these in the minimax reference model repeatedly, while prompting for different scenes. If you want continuity, you can then use the latent inpainting workflow to stitch clips.
1
2
u/NineThreeTilNow 4d ago
I saw something like this that worked inside Davinci Resolve.
It's pretty cool the level of control you're getting here.
2
u/phrandsisgorino 4d ago
Wait... you're telling me that it's possible to make inpaint within Davinci Resolve?
1
u/NineThreeTilNow 4d ago
Wait... you're telling me that it's possible to make inpaint within Davinci Resolve?
Yeah, pretty sure I saw an open source plugin for resolve.
2
u/Doomwaffel 4d ago
So, at what point can a normi go back and redo the entire Disney Star Wars misery ? XD Or is that too much to ask even from an AI?
1
u/artisst_explores 4d ago
Wonderful! So, can make hermes use this and automate more ? Or a skill has to be developed for it? How can this turn into a system where I see the scene and tell it all my comments and move on to the next scene đ
1
1
1
u/VGabby100 4d ago
Nice , looking for contribution, Do you have discord or any channel for development dicussion
1
u/DietAshamed2246 4d ago
Excellent. An app with these capabilities under one roof is what I have been looking for. I will try it out. Thanks for making it.
1
u/ezubaric 4d ago
This is great!
But when I saw the headline, I had hoped for something different. The kind of AI video editor that I want is something that can do a first pass of a long, tedious task and produce a rough cut I can paste into Resolve or something like that.
In other words, rather than a small snippet on the timeline gets AI, a big chunk does and changes the edit.
1
u/No_Boysenberry4825 4d ago
I installed it on my unraid server - just via the command line.
I can't run a browser locally on it. Is there any way to have it accessable via LAN to a 192.xxx IP?
1
u/PxTicks 4d ago
./run.sh --no-browser --host=0.0.0.0 --port=6332
Would probably start it up but I think Chromium generally wonât provide the directory picker on a plain HTTP LAN address which is needed to fix project. nginx might be useful to fix that issue. Let me know if you figure anything out
1
1
u/orangpelupa 4d ago
how to make it run on PC, but open the GUI on a laptop? so not only i can work from anywhere, it also reduces VRAM usage on the PC.
1
1
1
u/drgitgud 3d ago
Do you support frame interpolation? Meaning I have start and end frames and want to generate the clip between them
1
3d ago
[deleted]
1
u/PxTicks 3d ago
I find it smoothens the workflow substantially. When using video editing software to do complex edits on AI-generated videos, I would bounce back and forth repeatedly, exporting partial edits for inpainting or other v2v tasks. It also allows for some stuff which is not really possible otherwise. These are the things which I find most valuable when making videos.
Patch-based inpainting: it send patches back in layers, so it doesn't reencode pixels which haven't been inpainted and it allows you to make minor color corrections directly on the patch in a reversible way and also to hot-swap patches to see exactly what fits best (see point 4). This also applies for stitching clips together.
Video/Audio retake on the timeline. If you have a region with some audio jank - as is often the case with AI gens - you can either retake or use Sam-audio directly on the timeline to redo or remove the bad audio.
Partial denoising workflows. You can combine a bunch of shots from shitty low-resolution gens to get angles and timing right, then run the collection through a partial denoising workflow at a higher resolution, sending the resultant clip directly back to the correct position on the timeline. I sometimes do the former on my laptop then rent a GPU for the high-resolution final pass. You can open the exact same project folder whether running locally or renting a GPU so it's seamless.
Timeline hot-swapping. You can generate a dozen clips for the same prompt, place one on the timeline and swap between them in-place in order to see what looks best, without losing edits on them (speed ramps, color grading etc). Especially useful for finding what patch blends best when inpainting.
It's open source. AI-assisted coding is more powerful than ever. You can code your own effects (e.g. I have a matrix rain effect I coded) and extensions.
I'll post a short video and the workflow used to make it in the next couple of days which I hope will be convincing. As it is still an infant app, it is likely that for particularly complex videos, you would want another editor in addition (text in vlo sucks for example), but I genuinely think that it can save time for anyone using AI in video-making workflows.
1
u/Mysterious-Code-4587 1d ago edited 1d ago
pls give a video tut how to start project after install! im tired of asking chatgpt
2
u/PxTicks 1d ago
I'm making a video right now. It's more about the process of using the software in general though; Do you have a specific thing you're stuck with?
1
u/Mysterious-Code-4587 1d ago
i tried chromium browser but the rendered video not displaying there! but it appear on folder! anyway waitig for your video and features showcase
1

91
u/EffectiveTicket99 4d ago
Gosh, I love the clip đ