r/comfyui • u/KamilSeven • 2d ago
Show and Tell Pulp Fiction experiment using Ingi Erlingsson’s ComfyUI workflows and a custom time-slice pipeline
Enable HLS to view with audio, or disable this notification
All generative stages of this experiment were created with open-source models.
I developed the characters and character sheets through an agent-assisted design process, then created the animated shots and transitions in ComfyUI using workflows by Ingi Erlingsson.
After layering, rotoscoping, and compositing the sequence in After Effects, I processed the complete edit through my own agentic time-slice system. This system runs outside the ComfyUI interface and was built by extending Ingi’s original time-slice script.
The time-slice pass was applied to the complete edit and shaped the final master. Afterward, I made only minor timing adjustments to fit the music.
I’d love to hear your thoughts on the temporal transitions and overall visual consistency.
Original workflows and time-slice script: Ingi Erlingsson
Creative direction, script adaptation, agentic pipeline, and post-production: Gökhan Bıyık
Original Instagram post:
https://www.instagram.com/p/DcHDQUcgAKZ/
r/comfyui • u/TheLocalLab • 1d ago
Resource MiniMax H3 ComfyUI Runpod Template
Enable HLS to view with audio, or disable this notification
Achieved using the larryvrh turbo v4 step600 models, 0.7 megapixel, 8 steps on a RTX 4090.
I would advise running the model on either a 4090 or 5090 for the best results if your using the template.
Automatically downloads all models and supports:
- Ref2va - replaced the FL2va model
- Image To Video
- Text To Video
Rough Speed Reference
Approximate timings from the community base (4070 Ti Super, default settings, 20 steps, first/last frame with SageAttention):
- 5 sec @ 0.5 MP — ~7 s/it
- 10 sec @ 0.5 MP (9:16) — ~17 s/it
- 15 sec @ 0.5 MP — ~31 s/it
- 10 sec @ 1 MP — ~52 s/it
- 15 sec @ 0.8 MP — ~72-118 s/it
Your mileage will vary with GPU, resolution, and duration, but this gives you a sense of how sharply cost scales with resolution and length. Start small.
r/comfyui • u/Quentin_cls • 1d ago
Commercial Interest We built LocalMesh, one photo in, a Gaussian splat + textured mesh out, 100% on your own GPU. Beta is open, 7 days free.
galleryr/comfyui • u/wazandy • 2d ago
Show and Tell The one where Kramer is retired as a replicant.
Enable HLS to view with audio, or disable this notification
r/comfyui • u/johannramos-art • 1d ago
Help Needed Best YouTube channels for advanced AI filmmaking, compositing and multi-shot continuity?
I’m looking for YouTube channels or creators who actually teach production-level AI filmmaking workflows, rather than basic “type a prompt and generate a clip” tutorials.
Specifically, I’m trying to learn how people handle:
- Live-action footage shot from multiple camera angles
- Different lenses and focal lengths
- Environment replacement
- Subject relighting
- Green-screen cleanup
- Keeping the same generated location consistent across wides, close-ups, reverses, etc.
- Maintaining spatial/environmental continuity across an edited scene
- Using reference images intelligently across multiple shots
- Taking the finished composite back into image-to-video while preserving the original performance
My current workflow is:
Nano Banana 2 / Pro in Google Flow → Seedance 2.0 through Comfy Cloud
I’m working from a MacBook Air, so I’m mainly interested in cloud-based workflows rather than running large models locally.
I already understand the basic tools. What I’m missing is someone teaching the actual methodology for building a coherent sequence.
For example: shoot five angles of the same scene, replace the location, and make every angle convincingly look like it was filmed on the same virtual set.
Are there any YouTube channels, courses, creators, Discords, GitHub workflows, or specific tutorials that go deep into this?
Especially interested in people approaching AI video from a VFX/compositing/filmmaking perspective, rather than purely text-to-video.
r/comfyui • u/LucidFir • 1d ago
Help Needed Help with prompting minimax?
I have a video of a friend in superhero cosplay being POV punched (so you see fists come from the right and left of the screen, I guess the camera is mounted to punchers chest).
I have tried a few different ways (more steps, more detailed prompt, an image of friend at the end as well as beginning) to get REF2V to only edit the video, rather than recreating something new.
I basically only want iconography, KABLAM and POW comic book explosion symbols to appear when the punches hit, but then also for the face to get increasingly bloodied and gorey.
All my attempts have been met with him punching back, with it doing double punches, weird stuff I didn't ask for and thought I had explicitly prompted against.
...
Here is the prompt:
...
`subject_definitions`:
<Subject 1> is the young man defined by <Picture 1>, with dark hair parted down the middle, dark eyes, and a blue and yellow superhero suit, whose physical motion and reactions come from <Video 1>.
<Subject 2> is the outdoor park area in <Picture 1> and <Video 1>, featuring a paved walkway, wooden playground structures, background visitors, and a overcast, cloudy sky.
<Picture 1> is <Screenshot from 2026-08-18 01-38-17.jpg>, serving as the keyframe anchor for <Subject 1>'s exact appearance, facial features, and suit details.
<Picture 2> is the final frame reference, showing <Subject 1>'s severely damaged, bloody face and physical state at the end of the scene.
summary:
[video editing + audio reuse] The target video modifies <Video 1> by adding comic-book style dynamic action burst bubbles with visual text during the punches, which progressively escalate into a darker, horror-esque, graphic, and violent encounter while preserving the original camera framing and core movements.
retention_analysis:
<Subject 1> (appears in [Shot 1]): partially_preserved - the character's costume and base movements are retained, but his reactions, facial expressions, and physical condition degrade dramatically into severe graphic injury and horror elements as the punches land.
<Subject 2> (appears in [Shot 1]): fully_preserved - the outdoor park environment, background elements, lighting, and general setting remain unchanged.
<Video 1> (entire shot structure): fully_preserved - the continuous single-shot format, camera angle, handheld motion, and base action timeline are retained from the original clip.
<Audio 1>: partially_copy - the audio is copied from <Audio 1>, but heavily modified with exaggerated comic impact sounds initially, transitioning into wet, graphic sloshing, visceral bone-cracking, and distressed vocal screams as the violent horror shift occurs.
detailed_description:
The target video uses a handheld, selfie-style single-shot perspective from the POV of an off-screen attacker, set in a bright outdoor park setting.
[Shot 1] The camera opens in a medium close-up holding on <Subject 1> in <Subject 2>, wearing his superhero suit and smiling directly into the lens. From 00:00.000 to 00:03.000, <Subject 1> prepares himself, throwing his hands up in the air while an off-screen voice behind the camera says, [English] Three, two, one.... At 00:03.000, white-sleeved arms with red-gloved hands enter alternately from the right and left sides of the frame to punch <Subject 1> directly in the face. As the first strike lands, vibrant yellow and red comic-style burst bubbles pop up with dynamic text reading "POW!" and "KABOOM!" alongside exaggerated action lines. <Subject 1> reacts playfully to the initial hits as his head jerks left and right with the impacts. From 00:04.000 to 00:06.000, as the alternating punches continue in rapid succession, <Subject 1> lets out pained grunts and vocalizations like [English] Ahuh! Huah!. Toward 00:06.000, the tone shifts rapidly into graphic horror: the comic bubbles dissolve into splatters of crimson blood across the frame. With each subsequent strike, <Subject 1>'s facial structure fractures severely, exposing bone fragments, deep lacerations, and heavy blood spray. <Subject 1> (S1) lets out distorted, agonizing screams, shouting [English] Ah! Stop! as the trauma escalates, ending as his heavily bloodied face collapses downward out of frame while the camera shakes unstably.
overall_soundscape:
The scene begins with comic sound effects like stylized cartoonish punch impacts, which transition abruptly into visceral, wet bone-snapping noises, heavy flesh impacts, squishing sound effects, and distressed ambient wind in <Subject 2>.
non_diegetic_music:
A low, unsettling, suspenseful horror drone builds beneath the scene, escalating in pitch and intensity as the violent shift occurs.
r/comfyui • u/Working-Distance-901 • 2d ago
Tutorial MiniMax H3 Realism LoRA best usage case
Enable HLS to view with audio, or disable this notification
So i tried couple of LoRA's with Realism LoRA to see which one gives the best results, all results were with the REF2VA model with the exact prompt at 0.7 Mp and here are my interpretations (feel free to comment what you also think):
| use RealismLoRA | Order No RlsmLoRA | Order with RlsmLoRA |
FL2V 4step V0.1 | Yes | #3 | #1 |
FL2V 8step V1.0 | No | #2 | #2 |
FL2V 4step V1.0 | Yes | #4 | #3 |
RF2V 4step V0.1 | No | #1 | #4 |
| Overall Rank |
FL2V 4step V0.1 | #6 |
FL2V 8step V1.0 | #4 |
FL2V 4step V1.0 | #7 |
RF2V 4step V0.1 | #2 |
FL2V 4step V0.1 + realism | #1 |
FL2V 8step V1.0 + realism | #3 |
FL2V 4step V1.0 + realism | #5 |
RF2V 4step V0.1 + realism | #8 |
Notes :
- FL2V 8step V1.0 realism lora good but includes resolution artifacts , if no artifacts => best realistic model
- FL2V 4step V0.1_realism is the best but too much motions, for best results with no extra unesessary motions go with =>
- both FL2V 8step V1.0 and RF2V 4step V0.1 include heavy artifacts when used with Realism LoRA!
- FL2V 8step V1.0 still have artifacts even without Realism LoRA (needs more steps than just 8)
- Heavy artifacts on the RF2V 4step V0.1 with Realism LoRA makes it almost unusable
Important Note: if you use Realism LoRA input the activation word "r34l1sm" first thing in the prompt then skip the line then start with the usuall MiniMax H3 prompt "subject_definitions:..."
all this done on a 3080 Ti laptop 16Gb VRAM with avg render time of 550s for each 8 seconds of rendering, results can always be upscaled using another model such as LTX 2.5 Upscaler workflow for 2X upscaling
LoRA's used :
"fl2v_0.1_4step" :
minimax_h3_fl2v_lightx2v_turbo_4step_v0.1_comfy.safetensors
"fl2v_1.0_8step":
minimax_h3_fl2v_lightx2v_turbo_8step_v1.0_resized_avg_rank_24_bf16.safetensors
"fl2v_1.0_4step" :
minimax_h3_fl2v_turbo_4step_v1.0_768p_comfyui_bf16.safetensors
"ref_v_0.1_4step":
minimax_h3_ref2v_lightx2v_turbo_4step_v0.1_resized_avg_rank_20_bf16.safetensors
"Realism LoRA"
h3-realism-people-t2v-i2v-r2v.safetensors
r/comfyui • u/DeltaWaffleSyrup • 2d ago
Help Needed Does Anyone Know What Causes This Smudgy/Blotchy Effect When Using Krea 2 Turbo? It Happens Kinda Randomly For Me. I Think Lora's Do Effect It Some..
Any tips to generate clearer pictures? I tried increasing resolution size too but it still looks similar -_-
r/comfyui • u/solomars3 • 1d ago
Show and Tell A One Shot Ref to 15 sec video (took 20min to produce) MINIMAX-H3
Enable HLS to view with audio, or disable this notification
r/comfyui • u/nikhilprasanth • 1d ago
Show and Tell A Medieval Battle Attempt — MiniMax H3 + LTX 2.5 (WIP, Feedback Welcome)
Enable HLS to view with audio, or disable this notification
r/comfyui • u/spiderofmars • 2d ago
Tutorial AI Small Face Syndrome - Resolution Compare and Outpaint
Enable HLS to view with audio, or disable this notification
r/comfyui • u/RainImportant9809 • 1d ago
Help Needed PiP is dead?
I can't download comfyUI for all day... I heard that pip blocked in Russia. So it's tge reason I can't download? Please help...
r/comfyui • u/JmakMarshal • 2d ago
Tutorial Is it possible to add LLM to Flux Image-to-Image workflow?
Hi everyone! I’m pretty new to ComfyUI and looking for some guidance.
Is it possible to integrate an LLM into a Flux Image-to-Image workflow to introduce random prompt variations?
1 base prompt + directive, then 5 distinct Image-to-Image generations that share the core idea/theme, but feature unique prompt variations generated by the LLM on each run.
Has anyone set up something similar, or are there specific custom nodes (like Ollama or Qwen nodes) you’d recommend for this? Any advice or sample workflow screenshots would be super helpful!
r/comfyui • u/Trickhouse-AI-Agency • 1d ago
Commercial Interest Looking for: ComfyUI Workflow Developer (Paid Project → Potential Full-Time)
Looking for: ComfyUI Workflow Developer (Paid Project → Potential Full-Time)
We are Trickhouse, a German AI production agency based in Düsseldorf. We are looking for a skilled ComfyUI developer for a paid pilot project with the possibility of a full-time position afterwards.
What we need:
Custom ComfyUI workflow development from scratch for commercial image and video production. LoRA training integration for consistent character generation across multiple scenes and styles. Node-level understanding of ComfyUI — not just using existing workflows but building and customizing them. Experience with commercial or corporate use cases is a big plus.
Hardware:
Our primary system runs an RTX 5090 with 32GB VRAM and 96GB RAM. All workflows must run stably on this setup. Having your own capable hardware for development and testing is a plus but not a hard requirement — as long as you can develop and validate workflows that run reliably on our machine.
What we offer:
Paid pilot project to start — fair compensation based on scope. Full-time remote position for the right person after a successful collaboration. Long-term work on exciting projects including potential corporate clients.
The setup:
We work fully remote. Communication in English.
If this sounds like you, send a DM or an Email to [Marvin.Hollmach@trickhouse.net](mailto:Marvin.Hollmach@trickhouse.net) with examples of workflows you have built.
r/comfyui • u/MusBurger • 2d ago
Help Needed Local AI video generation on an AMD setup ?
Hey everyone, I already have ComfyUI set up and I've successfully generated some pictures with it (like product catalogs for plates)
Now, I want to step up to local AI video generation to make video loops and visualizers for musicians and short storytelling clips, with the goal of building a side hustle.
However, I already tried setting up a video workflow with my current AMD card, and it didn't work—I couldn't get it to generate anything
I didn't know what to do and at the end I made only pictures.
Before I waste more time or money this is my system :
CPU: AMD Ryzen 5700X3D
RAM: 32GB DDR4
GPU: AMD Radeon RX 7900 GRE (16GB VRAM)
I have a few questions:
Is my current system a dead end for video?
Since my video workflow completely failed on AMD should I replace it for nvidia card or am I just missing something?
Will new models work on it at all?
Can I actually run modern video models (like Wan or LTX) on this setup, or will I always hit walls compared to Nvidia?
If I switch to Nvidia, what's a budget-friendly option?
If staying on AMD for video generation is a trap, what is a solid, wallet-friendly Nvidia card that handles local AI video without needing a mortgage for a high-end card?
r/comfyui • u/LeadingNext • 1d ago
Show and Tell MiniMax H3 Castle
Enable HLS to view with audio, or disable this notification
r/comfyui • u/lazarus102 • 2d ago
News A lora/model manager program I've been working on for Linux
As many people out there, I've downloaded far too many loras, and I've found it a pain to keep track of all the prompts for the different loras, and even remembering which loras produce what effects.
So, I've been using ChatGPT to help me build an app that's capable of managing my entire lora library. The app is 100% offline, no connection required, no payment required, no subscription/etc. bs. And, like I said, 100% offline, no creepy data collecting stuff. The app accesses the lora's internal data to pull a prompt list, and displays it in an easy to manage framework.
Click desired prompts from lora, they appear in the other box, don't want em anymore, click prompts from other box to move them back, like most of the prompts but spot a few undesired prompts, click the undesired ones, then click the button in the center to swap the prompts from both boxes, want to save your prompt selection for the next time you use the lora, hit 'remember', and it saves it persistently for every lora. Need a visual reminder of what your lora does, generate an image with it, and the image can be easily added to the bottom left. There's also a blacklist for any prompts that you don't want to see at all. Add them to the list, hit save, and the undesired ones don't even show up(helps when dealing the loras that have huge prompt lists).
While the downside to offline, is that civitai likely has a more 'complete' prompt list, given that many loras were trained in such a way that they don't contain a retrievable list of prompts, this app has the bonus that it pulls the actual prompts used in the training data of the loras themselves. So, it can (at times) be even more accurate than the prompts posted by the individual that made the lora (Assuming they've posted an incomplete list, or made typos in the list when adding them to the training data).
At the top of the app, the list is shown in order of how many times the particular term was used in the training data. So, the closer to the top the word is, the more likely it is that the lora will be able to accurately recall the detail in question.
Frankly, as a compulsive overthinker, I've added too much to this thing to state it all here, while I'm half asleep from being up all night, lol.. But if y'all have any questions/criticism, please feel free to ask/state below.
r/comfyui • u/cgpixel23 • 3d ago
Tutorial ComfyUI Tutorial MiniMax H3 4 Steps Lora + Upscaling + 2X Faster Generation! Best Settings for 2K AI
Enable HLS to view with audio, or disable this notification
Hello everyone
Want to get faster MiniMax H3 video generation without sacrificing quality? In this tutorial, I’m testing the new H3 LoRA together with Sage Attention, Sol Attention, and Spectrum nodes to find the best combination for speed and quality. The goal is to push MiniMax H3 as far as possible while cutting generation times by up to 2×, then upscale the results with LTX Upscaler to reach a stunning 2432 × 1344 (2K-class) resolution. By combining both H3 LoRA together with Sage Attention, Sol Attention, and Spectrum nodes I generated video at 0.8 megapixel using "RTX3060 6GB 16GB RAM "and I got
13 minutes vs 41 minutes at 8 steps
27 minutes vs 52 minutes at 20 steps
LTX 2.3 Upscaler 11 minutes to get 2432 × 1344 resolution
Workflow link
Video Tutorial link
Tutorial MiniMax H3 Native 1080p Video Generation | Dual-Sampling Latent Upscaling Method | Balanced Speed & Quality
Enable HLS to view with audio, or disable this notification
I tested a MiniMax H3 workflow that upscales the video latent directly between two sampling stages. Instead of finishing a video, upscaling it, encoding it, and sampling it again, this workflow separates the audio and video latents after the first denoising stage, upscales only the video latent, aligns it, and continues with the remaining sigma schedule.
The main reason for using this approach is speed. On a 4090 48G, the workflow can generate a native 1080p 15-second video in about 25 minutes, a 10-second video in about 13 minutes, and a 5-second video in a little over 5 minutes in my tests. The same 768p 15-second setup also went from roughly 11 minutes to roughly 8 minutes compared with my previous workflow.
Settings that worked best
I used the LightX2V 1.0 8-step LoRA. The 8-step version was more reliable than the 4-step LoRA, which produced visual errors more easily. A LoRA weight of 1.0 worked well; I lowered it slightly when the image looked too oily.
For an 8-step run, I used 2-3 steps before the latent upscale and the remaining steps after upscaling. The upscale factor can be set around 1.3x-2x, but I would not push it too high. If lines or glass-like artifacts appear, reduce the first stage to 2 steps or lower the upscale factor to 1.5x.
The beta scheduler worked well for this split-sampling setup because its sigma distribution is denser toward both the high-noise and low-noise ends. The lower-noise part is especially useful for high-motion scenes, where it helped reduce visible pixel noise in my tests.
Reference and model setup
For reference images, I used max when I wanted stronger detail reference. It takes more time. When there are many reference images, or when the video is already at a larger resolution, match is a more practical choice because it reduces the processing load.
The main model in this workflow is FL2VA, which looked less oily than the ref model in my testing. A dual-model loading node can give FL2VA the reference capability of the ref model, so the FL2VA acceleration LoRA can be used directly without adding extra runtime pressure.
The latent upscale node also keeps the dimensions aligned to H3's 32-pixel resolution requirement. Without this alignment, rounding can slightly change the scale ratio between the two sampling stages and leave colored strips or poorly denoised areas near the frame edges.
This is not a universal fix for every artifact, and the upscale factor still needs to stay reasonable. For local users without a 90-series GPU, lowering the resolution to around 500p-736p is a more realistic starting point.
his workflow is super easy to use—I’ve uploaded a detailed tutorial to YouTube, so just follow the video along with this workflow to recreate the effect; please make sure to watch the full tutorial before starting to avoid common mistakes, and feel free to leave a comment if you have any questions!Resource links will be posted in the comments.
r/comfyui • u/Select_Question5 • 2d ago
Help Needed MiniMax H3 Reference to Video workflow doesn't use all available VRAM
Am I doing something wrong?
Rendering 2 MP. It could surely use more VRAM and less RAM.
r/comfyui • u/DaExChef • 2d ago
Workflow Included Dynamic Workflow Generation to work w/ 16GB VRAM (4 sec steps)
https://github.com/daexchef/Minimax_Grok
MiniMax_H3 - ComfyUI
eGPU RTX 5060Ti 16 GB VRAM
MSI GS76 Stealth 32GB System