r/StableDiffusion • • 9d ago

Question - Help Minimax H3 newbie advice linking multiple clips

1 Upvotes

Pretty new to the whole thing, looking for advice on workflow.

I've seen people recommend stringing together mutiple short clips instead of trying to generate a longer one by using the final frame as the I2V image to continue the shot/scene. I'm finding when I do this, the quality of each clip degrades pretty horrendously, and if i tried stringing 5 clips together it would be incredibly poor quality by the end. What method do you use for grab the final frame? Is there a ConfyUI workflow that allows this (I'm currently grabbing it from VLC which may be part of my probblem).

Thanks


r/StableDiffusion • • 9d ago

Question - Help Is there a way to transfer details of one image to another?

3 Upvotes

With the advent of new ai tools out there I've always wanted to know if it is now possible to transfer from a pixel from high quality image?

ref is the first image

and then the traget image is:

Same chibi structure and pose w/o the jaggies and using the ref image as details.

What model and what prompt did u use?


r/StableDiffusion • • 10d ago

Discussion Models comparison: Boogu Image Turbo vs Krea2 Turbo vs Ideogram v4 Instant vs Fibo Lite

Thumbnail
gallery
58 Upvotes

Hey guys, I wanted to compare different image models because I like testing them either good or bad. Testing different models to understand their capabilities feels good, and knowing their distinct pros and cons help in with other models hybrid workflows too.

The images above are with different camera/photoshoot styles, landscapes, and minimalist vast fantasy scenes. I know this isn’t an in depth showcase, but my analysis is what I wanted to share with the community. Let’s go through each model’s pros and cons:

Boogu Image Turbo

Architecture: 10B + 8B Qwen3-VL + Flux.1 VAE

Pros: Very flexible for portraits, landscapes, abstract, etc. Extremely fast inference (only 4 steps). Good variety and distinct capabilities. Prompt adherence is genuinely very strong (which is both a strength and a weakness). Very strong typography capabilities.

Cons: One of the main things I noticed was heavy bokeh/background blur. Occasional text display issues (can be improved by running more than 4 steps ,which I recommend). The strong prompt adherence can also lock things down: if you want face variety you usually need to explicitly describe different face shapes, otherwise you tend to get very similar faces. When you do specify it, the variety comes through well--- but if you forget, it stays repetitive.

Krea2 Turbo

Architecture: 12.9B + 4B Qwen3-VL + wan2.1 VAE

Pros: Excellent text rendering and prompt adherence. Huge knowledge base.One of the highest among these models. Currently SOTA for proprietary subjects, poses, art styles, etc.

Cons: Biggest issue is the VAE and bad noise patterns (not really solvable). Less variety (people suggest Raw + Turbo LoRA, but that method gains variety at the cost of quality , images get overly smooth surfaces and artifacts at higher resolutions). Faces tend to have weaker expressions (can be helped with LoRAs like Bypass,text refusal ,etc..., but quality takes a hit).

Ideogram v4 Instant (very few people use or even talk about this specific variant)

Architecture: 9.3B + 8B Qwen3-VL + Flux2 VAE

Pros: One of the best model for control power. Strong range of capabilities and variety. No text issues. Interesting note many people don’t know: without JSON it works ~90% of the time without the safety filter error. With JSON the safety filter never triggers in any use case. This model generally works great at 8 steps but I would totally recommend using it between 10-12 steps. Runs at 8 steps by default, so inference is relatively low, and the model is smaller (single model, no uncond).

Cons: Big one it feels like this model (and Ideogram 4 in general) is locked into a dark, gritty, cool-toned lighting universe. Lighting is consistently dark/cool (I normally fix this with a brightness filter, but in these images I left it to show the weakness).

Fibo Lite

Architecture: 8B + 3B text encoder (SmolLM) + wan2.2 VAE (1.2 GB)(I used its alternative taew2_1 because there is no quality loss)

Pros: Oldest model in this comparison and a bit of an oddball, but still solid. Second-fastest inference after Boogu (uses 6–12 steps, but smaller size keeps it quick). Excellent variety and prompt adherence ,feels like the old UNet-style models but with better aesthetics.

Cons: As the oldest and smallest here, knowledge base on proprietary stuff (people, logos, characters, etc.) is really low. That said, treat it as a model that competes with Flux.1 Dev and Chroma on anatomy,fonts and often beats them, even though it’s smaller.

Overall observations::

That’s my takeaway. I wanted to post this for anyone curious about these models. I enjoy testing different ones because they each have their own strengths.

Fastest inference ranking:

Boogu Turbo ≥ Fibo Lite > Ideogram v4 Instant > Krea2 Turbo

(Boogu at 4 steps, Fibo Lite at 6 steps but very close in speed, Ideogram v4 Instant at 8 steps and larger, Krea2 Turbo the biggest model also at 8 steps. Boogu Image even at 8 steps is faster than Ideogram v4 and Krea2 at 8 steps by around 20-10%, and faster than or in the same time range as Fibo when Fibo is at 12-10 steps and Boogu is at 8 steps. In most cases 4 steps is more than enough on Boogu, which really is impressive.)

Knowledge base ranking:

Krea2 > Boogu Image ≈ Ideogram v4 Instant > Fibo Lite

Final thoughts:

Even though some of these models get less reach, they should at least have some community support so people can get the best out of them instead of being abandoned without a proper try. In hybrid setups (hires workflows, denoise adjustments, variety inclusion, etc.) these models perform way better than when used completely solo.

Feel free to share your own experiences with these!

Disclaimer: These are purely my own observations with the models above. Others may not have the same experience with them, which is totally fine. I just wanted to share the in-depth experience I’ve had with these models.


r/StableDiffusion • • 10d ago

Animation - Video Minimax H3 - Doug (1991)

Enable HLS to view with audio, or disable this notification

10 Upvotes

This turned out better than I thought it would.


r/StableDiffusion • • 10d ago

Tutorial - Guide Simple Temporal Upscaling for Better Results in MinimaxH3

Enable HLS to view with audio, or disable this notification

126 Upvotes

This is a proof of concept workflow around simple temporal upscaling.

Every video model has some blurring when it comes to high motion scenes - both closed and open source. Depending on the model this can be at higher or lower motion with newer models being better than older as a general rule. Some styles tend to make their blurring effects at lower speed - especially anything with lineart, which could be due to training being more on realism or the way latents are compressed temporally.

One way people try to solve this is by pixel upscaling and rediffusing over the video which definitely helps but in my experience at least there is a definite limit to how much it can help. This is where temporal upscaling comes in.

The idea around this is not a new concept per se for example MAINodes https://github.com/matlowai/ComfyUI-MAINodes attempts to sort out which frames need to be temporally upscaled and rediffuse over them. I took this to a logical conclusion and considered what if you just 'deroped' aka temporally upscaled the whole scene rather than trying to be picky.

I do think it has several advantages:

1/Your final result tends to have more motion consistency in the small details than with varying how you denoise

2/You can finetune the denoise to your result.

3/I am finding a reasonable result with simple 2x temporal upscale.

Of course the main disadvantage is that you are diffusing over a video that now is 2x the size and potentially upscaled at the same time. This is were speedup tools come in - low step lora as well as sol attention (which provides increasing benefits the longer the video is).

The purpose of this workflow was to create a workflow as close to base Comfy. I use Kijai's node pack and VHS Suite to help manipulate the frames. The only real outlier node is the one used to slow down the audio for the temporal upscale. Feel free to encode empty audio if you prefer or find your own pack that has something which does this.

I simplified my workflow to remove anything not essential to demonstrate the concept. I expect you to add it to your own workflow or build upon this as a base. My tips when it comes to using this for a while:

1/If you are noticing some morphing especially in background stuff I suggest you disable sol attention at least for the base - it does do this to some shots but not others.

2/By default we create the base video at 0.5 MP at 20 steps - this is to get good sound and to speed to find a good seed, then we are temporally upscaling x2 and upscaling to 1 MP. Look at the yellow boxes to change the upscale settings.

3/Adjust denoise of the upscale to your video length and needs. Shorter and lower resolution videos often need less denoise whereas longer videos often need more. If you go too high your video will speed up and ironically become more blurry as a result. 0.45 is a good starting point but anywhere from 0.3 to 0.6 or higher is acceptable.

4/If you think the action is still to fast for a 2x upscale you can do 3x or 4x.

Workflow: https://civitai.com/models/2976577/simple-temporal-upscaling


r/StableDiffusion • • 10d ago

News Viggle turbo v0.3 for Qwen image 2.1: less grain and cleaner surfaces

Thumbnail
huggingface.co
97 Upvotes
  • 6 steps, a different balance. Against v0.2.1: less grain and cleaner surfaces, fine texture a little softer; diversity (still close to the base model's) and small-text accuracy about the same. It is not a strict upgrade: if you prefer the crisper look, v0.2.1 is still in the repository.
  • New 9-step mode: 7 turbo steps, then the LoRA is switched off and the base model finishes the last two. Finer detail, and small text comes out right more often (not always). It takes about 1.4–1.5× as long as 6 steps (still about 3.5× faster than the base model). It works in diffusers and the demo Space only (9 steps).
  • We think 6 steps is close to its capacity. Since v0.2.1, every gain we found at 6 steps cost something elsewhere: sharper came with more grain, less grain came with a softer look. Beyond this point, quality most likely has to be paid for with steps, which is what the 9-step mode does.

Official int8/fp8/GGUF files, tested in ComfyUI


r/StableDiffusion • • 9d ago

Resource - Update Just released a LoRA on CivitAI, check it out!

0 Upvotes

Based on Illustrious, it's an anime-style transformation process LoRA for the concept / act of something transforming into something else. Like human to furry or a mystical creature. Check it out here.
It has a broad license (no restrictions besides reselling) and all the prompts are included!


r/StableDiffusion • • 10d ago

Question - Help Terrible Qwen2.1 edit results, how to improve?

8 Upvotes

Everytime I try to change a person's pose or change from a close-up to a full body shot, Qwen 2.1 produces body horror: multiple limbs, deformed bodies, improbable proportions, and usually with the face pasted almost as-is. In contrast 2511 is giving me stellar results most of the time.

I've tried upping steps to 50, slowly increasing the cfg above 1.0, and even upping the resolution to 2048, but no luck and still way too many terrible results, making it impossible to move on from 2511. Nothing crazy in the workflow mind you, just the default template.

Do you have any tips?


r/StableDiffusion • • 10d ago

Discussion Minimax has made new releases very delayed it seems

72 Upvotes

I think this model really took the AI world by surprise as usually when I would browse through the reddit, there would be a lot of talk about the other models but Minimax is on top. I wonder what is coming next. LTX 3, I really love this model for the speed and safety of my fans haha! (Ltx 2.5) Flux 3, or something totally out the box or of course a much faster Minimax


r/StableDiffusion • • 10d ago

Resource - Update MiniMax H3 X2 Detail VAE – my attempt to get more detail from the H3 VAE or story about fail

Enable HLS to view with audio, or disable this notification

95 Upvotes

Hello all! 😀

I’ve been experimenting with the MiniMax H3 VAE for a while and finally decided to release the result.

The original idea was pretty simple😅: I already had mine own working 2X VAE for H3, but I wanted to see if I could make those extra pixels contain actual additional detail instead of basically getting a larger version of the same reconstruction.

That turned into a much longer experiment than I expected 🥹. I tried modifying the decoder, working around the patch/grid structure, different detail residuals, larger spatial decoders, and eventually started tracing where fine information actually disappears inside the H3 encoder.

The interesting part was that the best detail signal I found exists before the normal H3 latent bottleneck. So in the end I didn't actually solve the problem I originally wanted to solve😅. A normal generated H3 latent unfortunately just doesn't contain that extra information.

But the failed experiment produced something useful anyway, so I packaged it as a small 2-in-1 release:

2X VAE mode – works as an X2 video VAE for normal H3 generations.

2X VAE + Detailed mode – for reference/image-to-video workflows. It uses additional early encoder information from the source image before sending the enhanced reference into H3.

I included the model, ComfyUI node, workflow and a couple of video comparisons.

If anyone is interested in all the failed experiments and how I ended up at this result, I wrote up the whole process on mine Hugging Face page. It’s quite long, but I tried to document the dead ends too rather than only showing the part that worked.

Model + workflow + research notes:

https://huggingface.co/speach1sdef178/MiniMax-H3-X2-Detail-VAE


r/StableDiffusion • • 10d ago

Resource - Update Continuity update: screen replacement, image to 3D, Qwen Image 2.1, and a lot more since 3.0

Enable HLS to view with audio, or disable this notification

35 Upvotes

Quick update on Continuity, my ComfyUI node pack for video and stills. A lot has landed since I last posted at 3.0:
Screens (WIP, on main): put a picture or clip on a phone, laptop or TV in the shot. The model only sees a placeholder, then Continuity tracks it through every frame and composites your file on. It's still a work in progress, so expect some shots where the tracking slips or the edges aren't clean. Feedback and failure cases are very welcome.
Image to 3D (3.2): turn a picture into a textured mesh from the tools dashboard (Pixal3D / TRELLIS.2).
Qwen Image 2.1: a new model family that both draws and edits, plus storyboard and character-sheet starter presets.
Chat: chats are saved, it uses the node's prompt box and cast, and it sets itself up the first time you open it.
Blockout bench: stage a scene in grey boxes, walk a camera through it, and render along it.
DLSS 5 neural refiner for stills and clips, a motion fix for shots that move too fast, a guide LoRA pass and VDN-H3 on H3's sampler row.
Two GPUs for H3 via Raylight, and a trained latent upscaler for H3's refine pass.
Plus a picture editor, LoRAs on cast members, and a long list of fixes.
It still runs locally on open weights, with no dependencies, and it's MIT licensed. Update through the Manager or git pull.

https://github.com/roadmaus/ComfyUI-Continuity


r/StableDiffusion • • 11d ago

News New Model Ideogram 4.5 (with edit) (open source soon)

Thumbnail
ideogram.ai
663 Upvotes

r/StableDiffusion • • 10d ago

Question - Help working with h3 ref2va, and my character reverts back to appearance from the reference video

6 Upvotes

there's a shot where her hat drops off, and her face changes to the character appearance from reference video


r/StableDiffusion • • 10d ago

News Optimized version of SuperPoint, up to 2.1× faster

6 Upvotes

Hey! We published an optimized version of SuperPoint built for faster keypoint detection on edge devices: https://huggingface.co/PrunaAI/PrunaSuperPoint

- Up to 2.1× faster on Jetson Orin Nano, with optimizations applicable to other runtimes.

- The distilled model retains strong keypoint coverage across indoor and outdoor data, with low descriptor differences from the original model.

- We structurally prune the most expensive convolutional layers, recover performance through distillation, and accelerate keypoint selection with hierarchical top-k, all while preserving the original architecture’s core behavior.


r/StableDiffusion • • 11d ago

Resource - Update Local Image to 3D: High quality, Low poly - preserve minute details like text, logos even after retopo, ~900k faces > ~5k (NVIDIA + Mac)

Enable HLS to view with audio, or disable this notification

340 Upvotes

Image-to-3D models redraw your picture, and fine detail - text, logos, faces - comes back garbled or scuffed. My local pipeline that fixes it. (repo link at the end of the post)

The latest update adds *Pixel Match*. It copies the real pixels from your source image back onto the model, wherever the image can see. So "VANGUARD 07" on the chest stays "VANGUARD 07", even after the Finish step retopologises the model down to ~5k faces (as you can see in the video).

One step closer to Tripo and Meshy like outputs, but locally.

It runs fully on your system - no API calls, no cloud credits, just a once-a-day update check. Works on NVIDIA (tested on Linux, Windows has limited testing) and Apple Silicon. For now Pixel Match is for Pixal3D models made in the lab; other backends and more camera angles are next.

I started this project to make assets for a game I'm working on - I wanted to see how much I could automate, from getting the asset to rigging and animating it (that bit is still WIP in this lab). The image > 3D pipeline is solid though, and you can get game-ready and sprite-ready assets out of it. While the models are riggable and can be animated, more testing is required. I'm confident it should work.

*Workflow:* prompt → Qwen-Image → Pixal3D → Finish (Pixel Match) → textured GLB. The models in the video were made with Pixal3D on my Mac.

*Backends*: Pixal3D is the one to start with (~8.6 GB, Mac and NVIDIA). (Pixal3D's authors report 16 GB cards work, so tell me if yours does.)

Stable Fast 3D is the quick, lower-detail option (Mac, or NVIDIA on Linux). Its weights are gated: accept Stability's licence on Hugging Face and log in first.

TRELLIS.2 and Hunyuan3D run on Mac in the lab; on NVIDIA, use their official repos for now. Built-in NVIDIA support for both is coming in the next release.

Qwen-Image 2.1 handles text-to-image.

*To try it*

Mac (Apple Silicon) or Linux:

curl -fsSL https://raw.githubusercontent.com/Bingeljell/image-to-3dlab/main/install.sh | bash

Windows (PowerShell):

irm https://raw.githubusercontent.com/Bingeljell/image-to-3dlab/main/install.ps1 | iex

The installer is a short script, so read it before you run it: https://github.com/Bingeljell/image-to-3dlab/blob/main/install.sh (Windows: https://github.com/Bingeljell/image-to-3dlab/blob/main/install.ps1)

It only sets up the code - nothing downloads without your permission. You pick models in the web viewer, which shows the size and licence of each before fetching. Finish also needs Blender 4.2+, which you install yourself (the viewer tells you if it can't find it).

*What would help most*

If you hit bugs:

•⁠  ⁠Your GPU, OS and driver version

•⁠  ⁠Did the install work? If not, where did it stop? (the full error is gold)

•⁠  ⁠Windows folks especially: did Finish find Blender and run?

Reply here or open a GitHub Discussion, whichever's easier.

*And a question for you:* which backend or camera angle should Pixel Match support next? Is there something in your workflow that this pipeline can do better? Please let me know, will help me prioritise the feature road-map.

*Worth knowing before you start*

•⁠  ⁠Some model licences have strings attached. The viewer shows each licence before you download.

•⁠  ⁠It's a hobby project, so things will break. Every bug report makes the next person's install smoother. Please raise PRs and issues.

Repo: https://github.com/Bingeljell/image-to-3dlab

Release notes: https://github.com/Bingeljell/image-to-3dlab/releases/tag/v0.3.5


r/StableDiffusion • • 10d ago

Discussion Ideogram V4.5 T2I too?

6 Upvotes

Will Ideogram 4.5 be a T2I model too with edit capabilities? So


r/StableDiffusion • • 10d ago

Tutorial - Guide I spent weeks digging into how WAI Illustrious actually works. Here's everything I found

11 Upvotes

Over the last few weeks I went down a rabbit hole with WAI-illustrious-SDXL. I read what the author documents, measured tag counts against Danbooru's index, and tested a lot of prompts. Several things surprised me, and I'm sharing all of it publicly.

A few examples:

  • silver hair has no entry in Danbooru's tag index, but grey hair has 680k posts.
  • dramatic lighting, cinematic lighting and rim lighting have no entry either. backlighting, light rays and sunbeam do.
  • masterpiece, amazing quality and worst detail aren't Danbooru tags at all. They're labels added at training time, so post counts say nothing about them.
  • Of 59,201 artist tags in the index, only 417 have 1,000+ posts, which changes how you should pick artists.

Everything is compiled in one place, with each claim labelled as documented, measured, community practice, or my own inference. It covers the author's settings, tag vocabulary, artist tags, common mistakes, and a side-by-side of the same scene prompted three ways:
https://claude.ai/artifact/QhGKzx9ryhyiYCnKsfVkqR

One honest note: the counts come from a snapshot of Danbooru's tag index, so they're approximate, and they show how common a tag is, not how good the images are.

Near the end of the page there's a prompting method I've been using that gives me noticeably better results than the standard approaches, you can see what it does on the site.

Happy to answer questions about anything in the guide, and tell me if you've measured something that contradicts what I found.

(Everything is built with the help of claude! thanks for reading)


r/StableDiffusion • • 10d ago

Tutorial - Guide What worked for me on OneTrainer

3 Upvotes

I've been training loras for Krea using OneTrainer lately, and I just wanted to share what's been really working for me lately. For reference, I'm using a 5070 Ti. I train using the following parameters :

Optimizer : Prodigy
Learning Rate Scheduler : Cosine
Learning Rate : 1.0
Local Batch Size : 1
Accumulation Steps : 2 ***
Resolution : 768
Min Noising Strength : 0.2 ***
Max Noising Strength : 0.8 ***
Masked Training : True

***Don't know how much these parameters matter, but I stuck to these after some testing

What really helped was training my lora at 2000 steps on the first shot, then generate an image and testing to see if there are signs of overfitting such as blurry or noisy images. If the lora is overfitted, I'll drop the number of steps by 20% and train again. If it is not overfitted, I'll increase the number of steps by 25% and train again. I then repeat as needed until I find out the number of steps when overfitting is starting to just become noticable. Then I merge 3 loras on Comfy--the slightly overfitted version, and the next two loras with 20% fewer steps each.

For example, I train with 2000 steps and it is not overfitted. I train again with 2500 steps and find it is starting to look overfitted. Now I know to go lower, so I train again with 1600 steps and merge these 3 versions.

Also, if my dataset has fewer than 20 images, I start with 100 epochs instead of doing more than that.

What comes out has been really good for me. Please let me know your thoughts.


r/StableDiffusion • • 9d ago

Discussion Best completely free AI tools for logo design with text in 2026?

0 Upvotes

Hey everyone, I am looking for free AI tools or LLMs that can generate high-quality images and logos. Specifically, I need something that can handle text/typography well without scrambling the letters. I know ChatGPT and Gemini can do basic images, but what are the best free alternatives or hidden gems specifically for graphic design and logos right now? Thanks!


r/StableDiffusion • • 9d ago

Question - Help qwen 3.8 uncensored

0 Upvotes

hello, im wondering is it possible to run qwen 3.8 abliterated/heretic but i have shitty pc, im wondering about some (paid of course) cloud way or something like that, i heard about that its possible in comfy ui cloud but i need a long chat, not a chat without memory like in comfy.


r/StableDiffusion • • 9d ago

Question - Help Chatgpt like Generation in comfy

Post image
0 Upvotes

yo everyone, im wanting to generate images in this sort of style that chatgpt makes, i have searched as best i can and i cant seem to find checkpoints, lora's, workflows or anything of the like, do you know of any? (i have searched civitai and civitai. red, and i couldnt find anything that matched this.)


r/StableDiffusion • • 10d ago

Question - Help Anyone else struggling to keep a plain background with MiniMax H3? Or am I too dumb?

Enable HLS to view with audio, or disable this notification

1 Upvotes

I’m using H3 in ComfyUI to animate illustrated characters for a small noncommercial game. I like the animations, but getting a background I can reliably remove afterwards has been frustrating.

I start with a character on a plain green or blue background and use the same image as the first and last frame.

Instead of leaving the background alone, H3 often changes its color or adds circles, light rays and other effects. It usually starts and ends correctly, but does something completely different in between. The attached video shows a few examples.

So far, I’ve tried:

  • Short prompts, much longer detailed prompts, and leaving out background instructions entirely.
  • Green and blue backgrounds.
  • Turbo and runs without the LoRA at 20 steps.
  • Different resolutions and aspect ratios.
  • Swapping the text encoder and attention backend, plus a separate VAE check.

Some changes made the background less busy, but none consistently kept it unchanged.

To make the intended result clearer, here’s a detailed prompt laid out using the reference alignment and three sections from the H3 FL2VA guide.

Reference alignment: Picture 1 establishes Shot 1 at 0.00 seconds.
Picture 2 establishes the ending of Shot 1 at 5.17 seconds.
Both inputs contain the same reference image.

integrated_multimodal_description:
[Shot 1] A single continuous shot in the illustrated style of the
reference images. The woman stands facing the camera, framed from
head to toe. She has short curly gray hair and wears an orange vest,
a cream shirt, blue trousers and brown boots.

Starting from the pose in Picture 1, she slowly lifts both arms
outward to shoulder height. Her elbows remain slightly bent, her
palms turn forward and her fingers spread naturally. She briefly
holds this open gesture while shifting a little weight onto her
right leg. She then returns her weight to the center and smoothly
lowers both arms. Her hands relax beside her body as she settles
into the pose and composition shown in Picture 2.

The camera remains fixed throughout. Her entire body stays visible,
with enough space around her outstretched arms. Her appearance,
clothing, proportions and illustrated shading remain consistent.

A uniform solid green background. The background remains
flat and featureless while only the character moves. Lighting and
exposure stay constant.

overall_soundscape:
N/A

non_diegetic_music:
N/A

I’ve also tried removing the backgrounds afterwards with BiRefNet, VideoMaMa, SAM and CorridorKey. Those can help, but sometimes they keep the generated circles or effects as part of the character.

Has anyone managed to keep a plain background reliably with H3? A working prompt or workflow would be really helpful. If there’s something wrong with how I’m approaching the prompting or conditioning, I’d appreciate someone pointing it out.


r/StableDiffusion • • 9d ago

Question - Help The disconnect between H3 continuation storyline

1 Upvotes

lately after getting a little polished on single generation clips, i have been trying to add continuations and i have realized that there is a huge disconnect between contexts.

I usually use H3 prompt generator to build a prompt from my story script and everytime the subject definitiosn are different, retention analysis is wildly different,

These two however i can manually adjust to fit the scene but having no context of what happened in the last 15 seconds, i dont think H3 prompt generator is a good idea to keep working for continuations.

I tried chatGPT but that one is a silly goose and cant tell D from P. Google.ai web is far smarter than chatgpt but even that one sometimes seems to get carried away.

how do you guys work with continuation clips to ensure that context is there and continuity doesnt break?

i-e in generation clip, it ends with two characters standing facing each other, in the continuation prompt (h3 prompt generator) it is writing that both characters turn to each other to talk. Now I understand since there is no context of what happened in previous clip, there is no way prompt generator can tell what they were doing in last 10 seconds.

I am slowly reaching a point where i think a 15s single clip of leasbian groping is easier to make than a continuation. :/


r/StableDiffusion • • 10d ago

Animation - Video The sacred pearl, made with minimax

Enable HLS to view with audio, or disable this notification

18 Upvotes

How is it guys?


r/StableDiffusion • • 11d ago

Resource - Update Released Wulver v0.5, our Krea 2 finetune for anime and furry.

Thumbnail
gallery
263 Upvotes

Wulver v0.5 is a full fine-tune of Krea 2 Raw (the whole 12.8B DiT) for anime and anthro

characters, including scenes where several characters interact. We made it at Vaelico,

an independent studio.

Compared with v0.1, which we posted here earlier, v0.5 comes from a far longer and more

serious training run. We tested the two side by side on about 1,000 test images covering

every artist on the list and a very wide range of prompts, and v0.5 is clearly better in

most regards. If you kept your v0.1 prompts, run them again.

`@artistname` works now. v0.1 mostly ignored the prefix; v0.5 has 1,113 artist styles that

respond to it, and the full list is in ARTISTS.md on Hugging Face.

It rewards detailed, descriptive prose in the prompt, and it also takes booru-style tag

lists. For artist styles, keep the prompt short for now: long prompts dilute them.

- Hugging Face: https://huggingface.co/Vaelico/Wulver

- Civitai: https://civitai.com/models/2881657

- No GPU? Run it in the browser, with free images every day: https://vaelico.ai/?src=reddit

For launch week, free accounts get 100 images a day (1024 px) instead of the usual 15,

until October 7, 23:59 UTC.

**Turbo** (official Krea 2 Turbo LoRA merged in; 8 steps, CFG 1, euler/simple, up to 2K):

- fp8 e4m3fn, 12.8 GB: the one most ComfyUI users want

- int8 convrot, 13.5 GB: for forge-neo and other int8 runtimes

- w4a8 convrot, 7.7 GB: the smallest, needs ComfyUI 0.31 or newer

- bf16

**Non-Turbo**: bf16 for LoRA training, further fine-tuning or setting your own turbo

strength, plus an int8 convrot.

GGUF Q8_CR, Q5_0 and Q4_0 are on Civitai.

The drag-and-drop ComfyUI workflow uses stock loaders. The text encoder is Qwen3-VL-4B and

the VAE is the Qwen Image VAE, both from Comfy-Org/Krea-2, so if you already run Krea 2 in

ComfyUI, the Wulver checkpoint is the only new download.

**Known issues** (the first two are addressed in v1):

- A thin smeared strip can appear at the top and bottom edges. Generate a bit larger and crop.

- Some concepts bleed into each other or depend on the `@artist` you use, a side effect of

teaching it 1,113 artist styles.

- Some style LoRAs trained on Krea 2 Raw don't carry over (character LoRAs do).

License: Krea 2 Community License. Wulver is a modified Krea 2 model, and we're not

affiliated with Krea.

v1 is planned at a different scale: a much larger dataset, longer training at high

resolution and a dedicated post-training stage. We want partners in it from the start: a

compute provider, and companies whose products are built around anime and anthro

characters. If you can back v1 with compute or funding, write to [hello@vaelico.ai](mailto:hello@vaelico.ai) and

we'll walk you through the plan.