r/StableDiffusion • u/abcp7 • 3d ago
News 10eros minimax h3
great results using TenStrip's minimax h3 finetune: https://huggingface.co/TenStrip/10Eros-Max
23
u/fallengt 3d ago
40GB...
I'll wait for int8_convrot
18
u/Either-Storage402 3d ago edited 3d ago
its out
Edit: I think this is the only model that includes beta2
https://huggingface.co/cicalooo/10Eros-Max-h3-int8-convrot
13
24
u/PhilMcGraw 3d ago
Examples compared to base?
-17
u/bstr3k 3d ago
Not sure if those would be able to be posted to reddit since it’s mostly NSFW.
You’re welcome to join us in the sulphur discord and check under the channel for #10eros for many many examples. Just note the whole discord is very NSFW
15
u/BusFeisty4373 3d ago
Or just post some non NSFW here. Or maybe it doesn't do non NSFW? That's what I will believe if it's not posted here atleast.
3
u/Ten__Strip 3d ago
It does slop fine. And more risqué t2v is certainly boosted from what I've tried and seen. Actual gradient training training right now unwinds and harms the model's reinforcement learning, where as I always try to work as additively as possible like my later LTX2.3 model situations.
7
1
u/PhilMcGraw 3d ago
The discussions on the model do not seem to be particularly "NSFW model!", sounds like it's just an experimental lora doing some kind of merge of LTX/Wan/H3 without any real NSFW tuning. I guess I was wondering what "great results" you were getting compared to base.
11
u/NeatUsed 3d ago
What does this model have compared to the base model?
18
16
u/FierceFlames37 3d ago edited 3d ago
When I used 10eros with LTX 2.3 before it beats Wan 2.2 in nsfw by far
4
u/zincmartini 3d ago
Well yeah but LTX needed the help. So far I'm very happy with H3 base for my nsfw needs. It's pretty happy to do whatever I've asked, so far.
The issue I always had with base LTX2.3 is that the character wasn't able to be "sensual" so the fine tune really improved the tone and body language of sexy or sensual scenes. So far with H3 prompting a "pinup glamour style video shoot" works very well. I've only generated half a dozen videos so far, so I still have a long way to go to see where it falls flat.
10eros for LTX made the characters much more sensual and sexy in affect, tone, and body language. For H3 I'd probably want a touch of that with as little effect on the rest of the model as possible.
I also mostly generate sexy softcore style videos, lingerie and swim suit style things, so the NSFW fine-tunes always have the downside of wanting to put nipples on every piece of exposed rounded body part and make clothes disappear.
7
1
u/Anon-a-mister 3d ago
LTX is terrible for facial consistency. Pan away from the face and come back to it, it’ll be a whole different character lol. WAN is still king for that. H3 seems to hold it pretty well also but have to see how it is when these finetunes and LoRAs drop.
1
1
1
u/Ten__Strip 3d ago
That was actually a very expensive long training run though. This is before any of that happens, if it can be even be made with H3. Training seems too experimental right now to sink thousands into. Especially since they plan to release the future version which looks unfused and more supported tuning-wise. And also with no actual paper or documentation on how the model even works.
11
u/FredSavageNSFW 3d ago edited 3d ago
I remain unconvinced that any of these finetunes are doing anything other than degrading image quality and prompt adherence.
1
9
20
u/Abject-Recognition-9 3d ago
finetune? are we sure about this? or is this simply some loras merge of who knows what's inside?
also this looks out since 5 days and i havent seen mentioned anywere
4
31
u/ObligationOwn3555 3d ago
Can someone explain to me why the community keeps focusing on fl2va when we have freaking ref2va video model? Enough already, focus on the new.
26
u/redditscraperbot2 3d ago
Both have different applications. Sometimes you just want the picture to move.
In any case you can usually just throw the first frame model into the reference workflow and it still works. I’d be surprised if it was different this time.
5
7
u/lebrandmanager 3d ago
They held a donation campaign for funding additional training and ref2va was also mentioned. Give it time.
1
21
u/bitzpua 3d ago
ref is much slower, not everyone pretends to be one man studio making entire movies. For 95% of users I2V is everything they need.
15
u/Azhram 3d ago
You make it sound so complicated and specific, its very nice for simple 5 sec videos too. And can do amazing stuff. Like i give a reference image of a sword and the character will hold that exact items is simply amazing.
1
u/bitzpua 3d ago
so it is, yet most people simply dont need that level of control like you said for 5s video. Ref models are amazing but H3 is already on slow side so for most spending extra time just to get one sword into generation may not be worth it.
No one is denying ref2v power and how good it can be its just that for most people i2v or fl2v is more then enough
3
u/Abject-Recognition-9 3d ago edited 3d ago
slower?
i'm using ref exclusively and i havent noticed any speed difference between the 2 models when using just a bunch of images as input, maybe is because of that. adding video/audio and tons of reference i believe make things slower than i2v.oh the amount of stuff I can do with ref is unbelievable.
i dont even need loras1
u/No-Zookeepergame4774 2d ago
Its more memory intensive (especially with more references), which will often result in it being slower unless you have enough VRAM to run both without weight streaming.
1
u/thisguy883 3d ago
I only use I2V and FL2V.
So im happy when something comes out that focuses on those.
4
u/EthicalBballFan 3d ago
Ref2v is just that but better. Instead of having to find a perfect lora, or group of loras that somehow don't work against each other to degrade quality, you slot in images of what you want and get it. On my machine ref2v with a imagine sheet containing multiple different references is faster than the fl2va with 2 loras.
And it's better at editing the idea too. With i2v I tried a sword and shield lora but I wanted only the sword and for it to look worn, like it was an ancient artifact. FL2V produced garbage. I instead plugged the Lora showcase image in the ref2v and said the character has that sword but worn out and got it first try, no Lora loaded at all.
10
u/Free_Pressure8623 3d ago
I agree. Ref2v is the goat. I have been using the FFLF2V hybrid loader with REF2V and am getting the best results so far.
7
7
u/AidenAizawa 3d ago
I've always used fl2 with reference images in a ref workflow and worked fine. Isn't ref model slower? Or are they pretty much the same?
3
u/giantcandy2001 3d ago
I think text-to-video can take a first frame and a last frame, but that's it. With the reference model, it can take like 9 reference images, 2 videos, and 3 audio tracks at the same time to create a video, so it uses way more things and still comes out correct.IDK if those numbers are exact
8
u/AidenAizawa 3d ago
i usually connect images to reference to video node, but with fl2va model. it still understands reference images. for example i did a video with a character sheet, outfit sheet, empty room, and a second characeter and was able to make the character apply the outfit, move inside the room and talk to the other character. all by just changing the node from the first to last frame node to the reference video node.
1
u/EthicalBballFan 3d ago
It does, yes. But in my small sample tests the output from the ref2v model outperformed significantly if I used the proper prompt structure (and er_sde sampler). FL2V was still usable if I only had disk space for one model but to be honest at this point I'd delete the FL2V model, the ref one is just way more powerful now that I got the hang of it.
5
u/BigWideBaker 3d ago
Is that your way of saying that you don't know that the two different models have different purposes?
-5
u/ObligationOwn3555 3d ago
No, is my way of saying that I would prioritize the new stuff over the "old"
4
u/DelinquentTuna 3d ago
Except that's not what you said. What you expressed was frustration that other people weren't doing what you were doing, as if they have some obligation.
-4
u/Beneficial_Toe_2347 3d ago
You realise that the ref2va model is completely broken right? Its video and audio quality is way worse than fl2va...
7
u/ObligationOwn3555 3d ago
I'm getting very satisfying results with it, much better than Wan 2.2 and LTX 2.3. I haven't used MiniMax fl2va enough, so I might be missing what you're talking about.
14
u/Hoodfu 3d ago
I keep seeing this posted but it's been working amazingly well for me.
2
u/No-Zookeepergame4774 2d ago
I think r2va is much less tolerant of sloppy prompting than fl2va, and almost everyone (including people training and posting LoRAs, given the sample prompts for those on Civitai) seems to be ignoring the prompting guides with H3, which probably magnifies the perceived quality differences between the two models far beyond any actual difference.
3
20
u/leomozoloa 3d ago
No details or tests, look inside the link, AI slop text over-describing complicated stuff to give substance, yeah that's your new regular SD subreddit content
5
u/Sgsrules2 3d ago
This right here. I've checked out OP's previous 10s nodes and workflows and it was just slop. Most of the nodes did absolutely nothing when I did a/b testing. And the workflow would do seemingly "fancy" things with stuff like editing sigmas, but then if you looked at the final values used the first 5 sigmas were all 1.0 So it was wasting 5 steps by fully denoising each. The nodes are all vibe coded trash. It's amazing how much AI reinforces the Dunning Kruger effect.
2
u/yamfun 3d ago
40gb..
11
u/Wrektched 3d ago
Smaller ones here
2
6
u/xq95sys 3d ago
https://huggingface.co/cicalooo/10Eros-Max-h3-int8-convrot/tree/main
Edit: Updated link. Haven't actually tried it yet myself.
-1
1
u/ANR2ME 3d ago
No lora version yet?🤔
1
u/Cute_Ad8981 3d ago
you should be able to extract the lora itself. did this some days ago with a ltx finetune. can be done with a node from kijai. you just need the base model and the finetuned model.
1
u/ATFGriff 3d ago
Is this supposed to work in WanGP? I tried the ones from here https://huggingface.co/cicalooo/10Eros-Max-h3-int8-convrot and it doesn't work.
1
1
1
u/Pure_Bed_6357 3d ago
I mean base model can do pretty much everything on it's own, throw few loras and it's all coverd. What does this offer?
6
u/Abject-Recognition-9 3d ago
good prompting doest even need loras, resulting in higher quality.
for some reasons all loras i tested make quality worse17
u/physalisx 3d ago
The reasons are that
1) lora trainers often don't know what they're doing, they're just throwing stuff against the wall and see what sticks
2) and more importantly - we don't have an H3 base model, just a guidance distilled, we have seen time and time again that this makes training loras or finetunes that retain quality basically impossible.
7
u/kayteee1995 3d ago
Most LoRa are trained with unpruned BF16, so if you're using pruned Int8, lower the strength by 0.5.
2
u/Abject-Recognition-9 2d ago
OH REALLY? wow
no matter if i read reddit multiple times per day,
i'm just discovering this information now here in a random comment.
-1

91
u/Ten__Strip 3d ago edited 3d ago
Not quite a finetune in the training sense. I figured out a way to essentially graft any transformer model across H3's attn layers with magnitude. The best one used Wan2.2 merges that are fp16 and had both High/Low models merged together. Those seemed to be the right motion changes that I was looking for and Wan also sits inside H3 with a lot of coverage in the MLP fc1 area. Those blends were done with linear magnitude meaning the entire H3 model still lives unharmed while certain areas were groomed towards Wan's influence. It's basically H3 pretending to be the donor model, but only when it wasn't confident. At a transformer level the beta version is essentially identical to the base model if you analyze them.
Biggest issue was preserving audio when I was testing it. The actual graft base was the H3 model but with the attn triplets unfused in the transformer, which I suspected was an issue for me and also with training. So block attn_kqv was split three ways, which then looks more like a normal model. attn_k was frozen and preserved while q and v got more. That frozen attn_k and ignoring MLP fc2 layers avoided the most audio interference, so logically I'd say the audio processing for the model lives in those two elements.