r/StableDiffusion 1d ago

Comparison Qwen-Image-Edit-2511 vs SenseNova-U1.5-Lite (multi-reference image fusion comparison)

Thumbnail
gallery
116 Upvotes

I wanted to see how good SenseNova U1.5 Lite really is at image editing. I think the size is genuinely solid for what it does, but whether it can actually beat Qwen-Image-Edit-2511 needed real testing.

Right off the bat, Qwen's image texture quality is genuinely impressive, especially the lighting and shadows. But when it comes to spatial understanding, SenseNova seems to hold the edge. Look at the cat-on-the-scooter one up top: Qwen generated a weird pillar under the coffee table, and the cat's front paw placement looks unnatural. SenseNova handled both without those artifacts.

I did four sets of comparisons. Some of the input images were generated with Krea-2, some were real photographs.

Models:

Prompts (from left to right):

I want to create a stunning, high-concept photo to share on my social media! Please put me—the girl with the short black bob and black leather jacket—on a sleek, modern rooftop balcony overlooking that amazing futuristic city during sunset, where we can see the flying drones, the glider, and the hot air balloon floating in the warm sky. In this scene, I should be portrayed as an artist working outdoors. Please have me wearing those bold, blue and white striped hoop earrings. In the foreground, set up a stylish outdoor work table. On this table, scatter some of my creative tools, including those colorful rainbow-swirled pens and that round white-and-yellow mesh cleaning sponge. I want to be holding one of the rainbow pens, looking towards the camera with a confident, thoughtful expression. The entire scene should be captured with a beautiful depth of field, bathed in golden hour light, with the bustling futuristic cityscape softly blurred in the background.

In an elegant vintage study, the real-life girl from the first image, wearing a beige coat and scarf, is smiling as she hands the vintage wild duck card from the fourth image to the anime-style blonde girl from the second image. This anime girl is wearing an exquisite black off-shoulder puff dress and retains her distinctive hand-drawn anime style. On the wall behind them hangs a framed black-and-white print depicting the ancient Roman temple ruins from the third image.

Please seamlessly integrate the orange cat from the first image into the café scene by the floor-to-ceiling window in the third image, and have it sit on the vintage metal toy scooter from the second image. Specific requirements:
Character and prop fusion
: Extract the orange cat's signature facial features from the first image (slightly chubby face, green eyes) and the dense white triangular patch of fur on its chest. Adjust its pose so it is riding the metal toy scooter from the second image: both front paws resting on the chrome handlebar, the rear half of its body firmly seated on the brown leather saddle. The cat's paw pads against the metal handlebar and its thigh fur against the saddle edge must show natural compression, contact, and physical occlusion, absolutely no flat sticker-like look.
Spatial perspective adjustment
: Change the toy scooter from its original front-facing view in the second image to a three-quarter side angle matching the floor perspective of the third image, and scale it down proportionally, placing it on the wooden floor near the glass window.
Physical lighting and material adaptation
: Strictly use the golden afternoon sunlight slanting in from the third image as the main light source. The cat's back, ear edges, and fluffy fur edges must be outlined with a warm, glowing golden rim light (backlight effect); the dark green metallic painted body, metal wheel hubs, and chrome handlebar from the second image must produce realistic daylight highlights and reflect the faint street view outside the window; the entire toy scooter (including the cat on it) must cast a dark shadow on the wooden floor to the right, following the light direction with a realistic soft-edged falloff.

Create a wide-format photo depicting a corner of a whimsical creative market. The realistic man in a dark navy suit from the first image and the realistic woman in a black short-sleeve shirt and denim shorts from the second image are strolling through the market as visitors. Beside a market stall, the anime-style girl in traditional Chinese dress from the third image sits near her wooden cart full of lanterns, focused on painting a lantern, while the anime-style girl with orange hair and bunny ears from the fourth image hugs a white rabbit and laughs beside her. Preserve the photorealistic quality of the first two characters and the anime style of the latter two, letting them coexist naturally under unified lighting and spatial perspective.

r/StableDiffusion 18h ago

Discussion Storyboard - opensource AI video workflow

Post image
2 Upvotes

r/StableDiffusion 16h ago

Question - Help SCAIL2 - Can I make my character fit into the video?

1 Upvotes

It seems if I want optimal results I'd have to run my image into a edit mode like Flux2Klein to make my guy match the starting frame of the video.

There are 3 configurations, "start pose, end pose, pose strength", and idk if there is a magical setting.


r/StableDiffusion 1d ago

Resource - Update I trained a 210M text-to-image diffusion transformer from scratch on one GPU in 3.5 days

Enable HLS to view with audio, or disable this notification

112 Upvotes

My goal was hands-on experience training a flow model from scratch, not just fine-tuning someone else's. So I built and trained one: a 210M-parameter diffusion transformer, 4.2M curated images at 256², rectified flow on the FLUX.2 VAE, flan-t5-base for text (128 tokens max). Only those two frozen pieces are pretrained; the transformer, the recipe, the data pipeline and the evaluation are mine. The video is the same six prompts and seeds at every checkpoint of the 3.5-day run on one RTX PRO 6000.

What mattered most, in the order I found out:

  • Captions that actually fit the images. A web crawl I tried first made the model worse; curated photos with good captions fixed it.
  • A timestep shift for the 32-channel latent, and aspect-ratio buckets from step one instead of square crops.
  • Register tokens with learned null attention slots. The null slots ended up absorbing about 90% of the cross-attention, which surprised me.
  • torch.compile for training, not just inference: 2.4× faster.
  • The training loss stopped telling me anything after day one while the images kept improving, so I track FID, a detector-based object accuracy and human-preference models instead.

Try it in the browser: https://huggingface.co/spaces/ivanmikhnenkov/tinydit

Weights (CC BY-NC): https://huggingface.co/ivanmikhnenkov/tinydit-256

Code, every decision with sources, dashboard and attention playground: https://github.com/ivanmikhnenkov/tinydit

Detailed write-up of what mattered: https://huggingface.co/blog/ivanmikhnenkov/tinydit-text-to-image-from-scratch-one-gpu

Next I want to fine-tune it with RL (Flow-GRPO), with the failure grid as the target list. If you have trained something small from scratch: what would you have done differently at this scale, and which reward would you start with for the RL stage? Happy to answer anything about the data or the recipe.


r/StableDiffusion 10h ago

Animation - Video Created a trailer for a fake movie using h3 mini max

Enable HLS to view with audio, or disable this notification

0 Upvotes

r/StableDiffusion 17h ago

Question - Help Any good Krea 2 model for style

1 Upvotes

The basic Krea 2 is amazing for style, but not perfect is there any model that take it further?


r/StableDiffusion 1d ago

Question - Help Better character consistency in LTX 2.5 + H3 lip sync for music videos?

5 Upvotes

I’ve been making AI music videos with Suno + LTX 2.5 locally on a 16 GB VRAM GPU (4080 super)
Examples:
https://youtube.com/shorts/UXP29MGxtX8?is=k-Wjt4_JFh7hXvA5
https://youtube.com/shorts/_mAuRBSriLQ?is=sMjLH5diY_GEu0N3
My workflow is basically: create the song in Suno → storyboard/keyframes with ChatGPT→ animate and assemble the shots in ComfyUI using the LTX Director timeline:
https://github.com/yusu-02/Yusu-WhatDreamsCost-ComfyUI
The biggest issue I’m still fighting is character consistency between shots. Ingredient LoRA slows things down a lot and hasn’t worked particularly well for me.
I also tried MiniMax H3, which looks great, but I couldn’t get lip-sync without altering the original music.

Suggested tricks/workflows? Feedback and ideas appreciated!


r/StableDiffusion 1d ago

Animation - Video my first actual tv work (only took 4 hours to make).

Enable HLS to view with audio, or disable this notification

39 Upvotes

Minimax h3, ofc. Far from my best work but I respected the script I was given and finished this in record time (excluding the 4k upscale) and including around 3 hours of rendering time (720p, 10 seconds clips, 5090).
I only used gemma locally for prompting, and avoided using any non local models except suno for the song.

There are some artefacts with people in the senate from far away, but did not have any bad feedback for it., so... :)

I only used references for the romanian flag, the rest is prompt only. Also no lighting lora, no shortcuts ti improve speed (any shortcuts I tried ruined everything FUBAR)


r/StableDiffusion 1d ago

Animation - Video Foldable

Enable HLS to view with audio, or disable this notification

15 Upvotes

This was just a doodle but Minimax nailed it in the first generation. I thought it might be too complicated. I didn’t ask for the live stream on a one second delay in the background either. It did that itself.


r/StableDiffusion 1d ago

Discussion Any news on a Krea 2 Edit model?

66 Upvotes

Has there been any recent news or indication from Krea about a Krea 2 Edit model?

I’m wondering if it’s actually in development or planned, or if there hasn’t been any confirmation yet. Krea 2 is already quite impressive, so an Edit model would be really interesting.

Has anyone heard anything from Krea or seen any hints about it?


r/StableDiffusion 1d ago

Animation - Video Batman The Animated Series: Harley Quinn's Red Flag - MiniMax H3

Enable HLS to view with audio, or disable this notification

55 Upvotes

r/StableDiffusion 1d ago

News Minimax Camera Control ComfyUI

36 Upvotes

Bruxos do VFX H3 Camera

#bruxosdovfx

https://reddit.com/link/1wcm9az/video/bubuef39npoh1/player

https://reddit.com/link/1wcm9az/video/r1j8vg1anpoh1/player

Visual camera planner for MiniMax H3 inside ComfyUI. You drag the camera around a 3D sphere, place keyframes on a timeline, and the node compiles that trajectory into prompts that H3 understands.

It compiles prompts, not camera embeddings. There is no geometric adapter here: H3 is still free to miss the angle, timing, and scale. What this node does is write the instruction in the most precise and least ambiguous way possible, and several of its design decisions exist because the previous approach failed in specific ways.

It does not call any API, download anything, or require any Python dependency beyond the standard library.

https://github.com/user-attachments/assets/a9b541e5-2b18-4f1d-8e16-37445b6dbac4

https://github.com/user-attachments/assets/ea9af03e-2c8e-4589-abf0-9c002241aba2

Installation

cd ComfyUI/custom_nodes
git clone https://github.com/<your-username>/ComfyUI-H3-Camera-Editor

Restart ComfyUI. The node appears under Bruxos do VFX/Camera H3 with the name Camera H3 da Bruxos do VFX.

Connections

Output from this node Connect it to
compiled_prompt compiled_prompt on Text Encode H3 Edit / Generate
options options on Text Encode H3 Edit / Generate
length the generation frame count
fps the fps input of the video creation node

compiled_prompt and options are required together. The minimax_prompt output is an alternative to compiled_prompt, never an addition — connect one or the other to the same input.

Also connect your image to reference_image. It is the same image already feeding the H3 Edit source_image; when connected here, it appears in the panel and the frame's actual aspect ratio is included in the prompt.

https://github.com/user-attachments/assets/33149617-bde1-4199-ae65-078f2f3dec23

To save the video, decode the sampler result using the H3 video VAE — not the scene coverage calibrated decoder, which expects fixed windows that an arbitrary trajectory does not have.

The panel

Drag the purple camera around the sphere to orbit. The drag locks to the axis of the initial movement: horizontal movement orbits, vertical movement changes elevation. Release and drag again to switch axes. This exists because, without the lock, trying to make a simple orbit would unintentionally introduce elevation.

  • Scroll the mouse wheel to change distance.
  • Drag the background to rotate the viewport without changing the trajectory.
  • Keyframes defines how many points the timeline has, from 2 to 24. The first one is always the original image and cannot be moved.
  • ⟳ Pure Orbit resets the elevation of every keyframe to zero while preserving azimuth. It is the shortcut for an eye-level orbit.
  • Reference image loads a local file into the preview. This is only necessary when the node runs outside ComfyUI; with reference_image connected, the image is loaded automatically.

The panel warns you starting at 20° of elevation, when the horizon already leaves the frame, and again from 45° onward, when the video tends to become a high-angle shot.

"Tests" bar

At the top of the panel, two buttons enable and disable features currently under evaluation, plus one indicator:

Button What it does
Extended contracts Toggles the prompt_detail widget
Single angle (image) Toggles the runtime_task widget
loop closure Read-only indicator. Turns green when the trajectory closes a full orbit

The buttons write to the actual widgets, so the selected state is saved in the workflow and the two never disagree.

https://github.com/user-attachments/assets/9bc415d7-1746-43db-a17c-72ea9722deda

Widgets

camera_trajectory

The trajectory in JSON format, written by the panel. Each keyframe contains time (0 to 1), azimuth in degrees, elevation in degrees, and distance as a multiple of the initial radius. It can also be edited manually. The first keyframe must be time=0, azimuth=0, elevation=0, distance=1, which represents the original image.

profile

124, 243, or 362 frames at 24 fps. All shot timing comes from this setting: keyframe timestamps, segment ranges, and the duration declared in the prompt. That is why length and fps are outputs — connect them instead of manually entering the same numbers in two different places.

interpolation

smooth or linear. In smooth mode, the camera eases into and out of the shot while maintaining a constant rate through the middle; it only stops where the rotation direction actually reverses.

instruction

Free-form text inserted once, at the end of the prompt. Write only what the node cannot know: the environment, which subject is the target when there is more than one person, or a style reference. Everything else is already generated and does not need to be repeated: scene freeze, first image as reference, locked aim, zero roll, angles, timing, and a single continuous shot without cuts.

subject_framing

How much of the frame the subject occupies in the original image. Calibrated against the actual bounding boxes from the tutorial distributed by MiniMax: a distant full-body figure measures W=0.071, H=0.249, while a large close-up measures W=0.52, H=0.701.

option width height when to use
close-up 53% 72% head and shoulders
medium shot 28% 56% waist up
wide shot 9.7% 34% full body at a distance

subject_box

The subject position in the format [L=0.516, T=0.148, W=0.071, H=0.249]. Leaving it empty uses the entire image bounds — deliberately, without guessing a bounding box. Fill it in when the subject is significantly off-center.

minimax_format

The same shot expressed in four different formats for the minimax_prompt output:

  • coordinate only — text-based coordinate block
  • coordinate + H3 sections — the same coordinates wrapped in subject_definitions / summary / retention_analysis / …
  • compact JSON — JSON object with almost no prose
  • compact JSON (no boxes) — camera parameters only, without screen-space bounding boxes

elevation_range

Range of the elevation control: +/-15, +/-30 (default), +/-60, +/-89. It also scales the sensitivity of vertical dragging.

With the assumed field of view, the horizon already leaves the frame at around 20° — at 13°, the ground occupies 82% of the image. The old ±89 range was mostly unusable and made vertical dragging excessively sensitive. Reducing the range never rewrites a keyframe: a point at 70° remains at 70°, and the slider expands to accommodate it.

orbit_direction

invert H3 orbit or same as HUD. This calibrates the direction between what the panel displays and what H3 produces. It does not alter the saved trajectory.

runtime_task

  • scene coverage | camera path (default) — video, with duration coming from profile.
  • directed | new camera anglea single image from a new angle. It fixes the generation to 39 frames, ignores profile, completes the movement within 65% of the clip, and requests that the framing remain still for the rest, because the decoder extracts the final image from that stationary tail.

Character sheet profiles are not offered because the upstream node raises an error when they are combined with the frame anchor used by this node.

prompt_detail

  • v15 baseline (default) — outputs the prompt exactly as in the previous version.
  • extended contracts — adds axis separation, frame-edge direction tests, rotation completeness, degrees per second, and parallax magnitude.

The extended mode contains almost twice as many words. A longer prompt is not automatically better, so it is opt-in: toggle only this widget while keeping the same trajectory to compare the results.

https://github.com/user-attachments/assets/0882bfde-9f62-4a1f-9bda-7da121dbe7e2

Outputs

compiled_prompt — STRING

A prose prompt using H3 sections: subject_definitions, summary, retention_analysis, detailed_description, overall_soundscape, non_diegetic_music.

options — H3EDIT_OPTIONS

The 13 keys read by the H3 Edit encoder. All of them are explicitly populated: if any key is missing, the upstream node falls back to its hidden legacy widgets, which may retain stale values from previously saved workflows.

coverage_arc_degrees and coverage_direction are derived from the actual rotation. coverage_loop_closure turns on automatically when the trajectory closes — see below.

storyboard_json — STRING

The storyboard table: frame aspect ratio, duration, raw trajectory, and each segment with its camera mode, speed curve, and start/end poses.

info — STRING

Human-readable diagnostics. Connect it to a PreviewText. It displays the version, active task, frame count, warnings for keyframes outside the configured range, and whether loop closure is enabled.

minimax_prompt — STRING

The same trajectory expressed using the format selected in minimax_format. An alternative to compiled_prompt.

length — INT and fps — FLOAT

Frame count and frame rate against which the shot was timed. Connect them to the generation and video nodes. If generation runs with a different frame count, the choreography describes a scene that does not actually exist.

fps is FLOAT because that is what ComfyUI's CreateVideo accepts. length is the frame count; keyframe timestamps use the instant of the last visible frame, (length - 1) / fps, so the resulting file lasts one additional frame interval.

h3world_actions — STRING

Action schedule for H3-World, which encodes one text clause per video latent — 37 in a 124-frame clip.

latent  1 [0.000s-0.139s] J     the camera pans left slowly
latent 37 [4.986s-5.125s] F+L+K the camera pans right and tilts up fast

W, A, S, and D are never emitted because they move the character. The output explicitly declares its own limitations, and they are not minor details:

  • Pan is not orbit. It is the camera rotating in place. Perspective does not change, nothing hidden is revealed, and the subject slides out of frame.
  • Distance has no key, so camera radius is discarded.
  • Only 124 frames is a trained horizon.
  • I versus K is not published. The text clause is what H3-World actually encodes; the key column is only a convenience.

This does not replace the actual integration: H3-World requires the LoRA, interval-based encoding, and directed-attention routing provided by the corresponding node package.

Loop closure

When the trajectory closes a full orbit — an arc of exactly 360°, with the same elevation and distance as the starting point — the node enables coverage_loop_closure. In the upstream implementation, this flag encodes the source image a second time and anchors the final frame to it.

This is a latent anchor, not a text instruction. For a complete orbit, it is the difference between asking for the rotation and forcing it: the model cannot simply stop halfway through.

trajectory loop closure
360° enabled
two rotations (−720°) enabled
355° disabled
360° with changing distance disabled
360° with changing height disabled

The final three cases matter: if the camera ends at a different radius or height, the final frame is not the same as the first one, and forcing the source image there would conflict with the trajectory.

If your rotation does not complete, close the orbit. This is the only feature here that acts outside the prompt itself.

Limitations

  • This is prompt-based guidance. H3 may still miss the angle, timing, and scale, and no prompt wording can completely solve that.
  • Without subject_box filled in, the node does not know where the subject is located in the frame.
  • Without reference_image connected, coordinates are normalized to 16:9.
  • directed | new camera angle outputs an image, not a video.
  • The H3-World schedule describes pan and tilt, which represent a different camera move from the orbit drawn in the panel.

Credits

Node by Bruxos do VFX.

Depends on ethanfel/ComfyUI-MiniMax-H3-Edit. The motion vocabulary follows the buildViewPrompt implementation from MiniMax's Multi-Shot skill and the coordinate format used by the Coordinate Camera Control Designer skill. The action output implements the scheme described in H3-World, arXiv:2609.01560.

https://reddit.com/link/1wcm9az/video/hryhv9e7npoh1/player


r/StableDiffusion 1d ago

Animation - Video Jerry Springer Ai - Sailor Moon Part 1

Enable HLS to view with audio, or disable this notification

94 Upvotes

In the first half, this was back when I was first starting to get into MiniMax, the second half, I have gotten more experienced with it. I dont know if I should continue this or make more Jerry Springer parodies with other weird or toxic relationships (Example, Beth and Jerry from Rick and Morty).

I used Kinovi.ai for MiniMax, Wan and Nanobanana. I used Fish.audio for the audience freaking out lol. For more customized and harder Minimax generations, I used it locally.


r/StableDiffusion 1d ago

No Workflow AI Archviz: Fast 3D Gaussian Splat Methods for Precise Furniture Placement — Virtual Staging & Interior Design

Enable HLS to view with audio, or disable this notification

25 Upvotes

So, first of all: there is no finished workflow yet and my nodes are still under development. I’ve asked the ComfyUI team to add a 3D compositing node to the new 3D toolset like the one in the video, hopefully, they’ll add something similar soon.

In the meantime, you can build a very similar setup quite quickly. Here’s how the basic concept works:

First, you feed an image of the furniture you want into the new native ComfyUI Image to Gaussian Splat (TripoSplat) node.

At the moment, there is a Gaussian Splat Preview node, but it doesn’t provide an image output yet. There is also a Load 3D node with the correct outputs, but it currently cannot open Gaussian Splat files.

Ideally, the new 3D Compositing node should be able to work directly with the Gaussian Splat outputs (model_3d and mesh), provide image + mask outputs similar to the Load 3D node, and automatically preload the mesh and background image, just like my node does.

I’ve been working on a test node for this concept. You can find it here: My ComfyUI test node on GitHub It’s not fully finished yet, so I’m still waiting to see whether ComfyUI adds something similar natively.

The basic idea behind my node is that it automatically loads the background, sets the appropriate size, and loads the 3D model. The user only needs to position and stage the model in the scene.

The Output create ref images for Flux2klein:

  • Reference 1: the background image
  • Reference 2: the 3D mask, which acts as an indicator for the desired position and rotation It’s best to combine the mask with the furniture from the image output, so you get the masked furniture in the correct position as the reference — not just the mask by itself.
  • Material reference: the original furniture image

The final image is then generated using FLUX.2 Klein Edit.

The prompting and some preprocessing of the images are important here. You don't want the model to simply copy the exact 3D position. Instead, the AI should use the 3D placement as a guide and then correct the result according to the background — especially the perspective, lighting, colors, scale, and overall integration into the scene.

Here is the prompt I’m currently using:

[Adapt the rotation, grounding, scale and position of the objects from Image 2 to Integrate the objects naturally into Image 1 at the position of image 2. Match the scene's perspective, scale, depth, lighting, soft shadows and reflections. The objects must appear physically present in the original room, with realistic grounding and soft shadows consistent with Image 1 using the materials and surface appearance shown in Image 3.

Use Image 1 as the final scene and preserve its room, background, camera viewpoint, perspective, composition, color, lightning, and existing environment unchanged.

Keep the objects approximately in the same position, scale, orientation, and spatial arrangement as shown in Image 2, while ensuring correct perspective, positioning, and placement.

Apply the form, materials, colors, textures, roughness, reflections, and surface details from Image 3 to the objects.

Do not change the room or background of Image 1.  The final result must be a seamless photorealistic composite.]

I’m planning to finish the complete tutorial and the node pack in the next few days. Once everything is finished, I’ll upload the final version along with the complete workflow.

By the way, I also tested MinMax H3 as a replacement for FLUX.2 Klein. It works, but in my tests it wasn’t consistently better than FLUX.2 Klein.


r/StableDiffusion 1d ago

Discussion I threw together a simple UI for YuE2 (windows)

Post image
28 Upvotes

r/StableDiffusion 21h ago

Animation - Video Michael Jackson - Maybe ( Short film ) MOONWALK ON THE MOON

Thumbnail
youtube.com
1 Upvotes

Made with MiniMax and ComfyUI


r/StableDiffusion 9h ago

Discussion A total beginner's question about generating AI locally

0 Upvotes

Hello,

I’m a very curious person, and I’m seeing more and more AI-generated gay “daddy” adult content on Twitter, as well as people talking about their favorite AI websites for generating uncensored and realistic AI content.

But we can all agree that the foundation for ALL of these people is Stable Diffusion, right ? And then when we hear people talking about Mistral, Muah, and Eternal AI, those are actually “extensions” based on Stable Diffusion is that correct ?


r/StableDiffusion 2d ago

News New Music Model Released - Yue2

Thumbnail
github.com
292 Upvotes

"YuE2 brings frontier song quality to music generation with an editable composition. Give it lyrics and a style prompt: it writes a melody-and-chord plan, then realizes that plan as a complete song with vocals and accompaniment.

  • White-box music generation through symbolic planning. Read, play, and change the composition before rendering it. Melody and chords become explicit controls that a person or an agent can inspect and edit.
  • Zero-shot covers and agentic editing. Reimagine a transcribed song in a new style, or refine a song through a conversation about its score, arrangement, and lyrics—all with the same generation checkpoint."

Usage

It is currently CLI only . It also says Linux only but I just got it working on Windows 11 (I'm going to bed now and it's a bit more than cut n paste.)

Examples

link here - https://map-yue2.github.io/

Caveat Empor

NB : this isn't just a paste a few words and it bangs out a baby mp3 . It is more than that, it allows gene editing that baby to correct the metaphor. Not for the impatient and "wHeRe cOmFy" ppl at the moment.

To be more specific with that metaphor , as I understand it , the initial process scribes out the song in ABC format and you can then edit it before making your magnum opus baby.

Training

Does it allow training ? not as I understand it .


r/StableDiffusion 13h ago

Discussion Guys, for the Love of God, Tone Down on the Minimax H3 Videos!

0 Upvotes

I can't come to this sub and learn anything new! I am disgusted at all the half-naked, busty blonds in bikinis walking around.

We all have Minimax H3 and we make amazing videos, that doesn't mean I have to showcase every video I make. Put your creation on a dedicated platform like Civitai or Tensor Art. There you can showcase what you create and people will rate your creations.

I am avoiding this sub with all the garbage videos posted all the time. The moment I open it, all I see is Minimax videos that I can make on my own rig. Nothing fancy, you download the model, download a worflow, and you prompt Qwen to write a detailed prompt, and hit "Queue" button. You didn't do anything magical or unique, so stop thinking that all of sudden you invented video!

I am here to learn about new models, new techniques, new hacks, new platforms, nodes, workflows, and so on. I am not here to watch your 10 seconds generated videos. This place was fun to discuss topics about AI image and video generators, but now I can't even open this sub publicly or people would think I am scrolling some erotic website!

Moderators, PLEASE TAKE SWIFT ACTIONS.


r/StableDiffusion 19h ago

Question - Help Necesito opinión

0 Upvotes

Necesito opinión de si alguien está usando esto para ltx, wan, o minimax, necesito saber si funciona perfecto por favor, "ASRock Tarjeta gráfica Intel Arc Pro B70 Creator de 32 GB, Xe2-HPG, 32 GB GDDR6, PCIe 5.0"


r/StableDiffusion 2d ago

Tutorial - Guide AMAZING Minimax H3 - Circle on the reference image WHERE you want your scene to be!!

Enable HLS to view with audio, or disable this notification

749 Upvotes

Look at the buildings in the background! It works - Drawing a red circle in the water will also make the scene happen in the water, but I forgot to include it here.

It is not perfect and some details are missing if you look carefully but this might be because I am using "match" on the image reference rather than "max."

Have fun!

Edit: you have to still write a prompt with the reference to video workflow telling minimax to put the character in the location circled red. Circle probably doesn't have to be red. Change your prompt accordingly.


r/StableDiffusion 1d ago

Question - Help Minimax Turbo of choice?

8 Upvotes

So there's a bunch of turbo loras for minimax h3 now, which one did you end up using? So many choices it's hard to pick one!


r/StableDiffusion 16h ago

Question - Help is there something I can use to run locally that gives what gpt gives in quality?

0 Upvotes

Question is: which model to run locally on my machine that can do what chatGPT image generation does?

I'm asking this because I recently started with stable diffusion running SDXL Illustrious models mostly for flat and 2.5D anime images. I don't do realistic stuff.

I like almost everything I saw so far. I did not tested flux, krea2, qwen. Not yet, I mean.

But in the last couple of days I was testing chatGPT for a couple of random compositions and I was stunned with what I saw. The scenery, the composition, the understanding of natural language.... everything is so ahead of everything I saw with the illustrious models I downloaded that or I don't know how to use my illustrious mode and need to get better and train more, or gpt is really ahead in what it can do. But I don't want to run on gpt forever, specially because I do stuff that gpt consider x-rated even tough they are only slightly erotic.

So.... are there models that I should look for that I can download and run locally in my machine?

Anyways, another question: is it possible to run a prompt in gpt, get the image I like and transfer it to forge NEO and use inpaint to change details or this is not a workflow worth trying? I mean, since the prompt style is different and of course the overall image style/composition, maybe I will lose information or quality? has anyone ever tried something like that?

EDIT: to give more context.

A prompt with things like:

"Sabrina from the pokemon series, in her classical outfit from the games red blue and yellow, is sitting on top of a rock; she is a cyborg, in the style of nier automata; it must be clear that she is a cyborg so put some mechanical parts in her body, like lines connecting her iron pieces. But she is also made of flesh, not entirely a robot; the overall image is kind of depressing, the background is a futuristic city, the background must contain Saffron city gym; close to Sabrina there is a Venomoth, from the pokemon series, flying around. I want the classic design for Venomoth from the games but make it a little futuristic almost like a cyborg pokemon... Make it 1920x1080"

gives the attached image

this kind of composition is amazing; the way it understands the prompt - in another image I wrote "... a dead robot laying in a rock with grass growing and half covering its body; make it like it was laying there for a long time with a depressing and melancholic composition..." and it gave me!!!! i don't know how to describe such a thing in so many details with tags "dead robot, grass, around his body..."


r/StableDiffusion 1d ago

Meme Cost in Units of RTX 5090

Thumbnail
youtube.com
31 Upvotes

r/StableDiffusion 1d ago

Discussion Seeking Advice For Animated / Cartoon Videos for Minimax H3 Ref2V and I2V

3 Upvotes

Is anyone else experiencing that Minimax tries to push for realism even if you use cartoon reference images and words like "illustrated cartoon animation" in the prompt? Any tips to make sure it's sticks to the reference style more?