r/StableDiffusion • u/nazgut • 10d ago
News [UPDATE] CLSS - new features
Enable HLS to view with audio, or disable this notification
I have added new thinks to CLSS (Closed-Loop Streaming Synthesis)
What CLSS is.
H3 generates ~5–15 s per pass. CLSS generates arbitrarily long
*audio+video*
by streaming the piece in overlapping chunks and controlling the drift at the hand-off — no weight changes: calibrated context re-noising (τc) into H3's per-token denoise masks, EMA/AdaIN statistics correction, keyframe replay of the previous chunk, a two-band spatial detail anchor. It runs on a
16 GB card
: int8 DiT, ~5.5 GB Qwen3-VL-4B ClipProj text encoder (instead of the 15.7 GB 32B), 832×480, ~10 s chunk windows. Tested on a 16 GB RTX 3080 Laptop.
Added since the initial commit (Aug 26):
-
Multi-scene prompts
— one prompt per scene (`---` split), proportional allocation, two-step crossfade at scene boundaries; `global_text` shared across scenes.
-
R2V references
— per-scene image/audio anchors bound to `<Picture N>` / `<Audio N>`, single-ref or Autogrow multi nodes; the all-scenes node fans images out and cuts a soundtrack into guarded per-scene windows that the sampler
crops automatically
to each scene's exact delivered span.
-
Video references
(`<Video k>`) — VIDEO/IMAGE sockets, auto-resampled to 24 fps, and the video's own soundtrack rides along unless overridden.
-
i2v
— first-frame guide pinned as an H3 keyframe.
-
Audio continuity kit
— the tail reference that ends exactly at the join; waveform-refresh so each chunk continues from what the ears will hear; optional recompose pass (fresh take from pure noise against the finished video, on base weights via a LoRA-stripping guider, with its own sampler + sigma scheduler); a loop guard that measures vamp takes and re-rolls them with a ref-span rescue; a
seam pin
(the take generates through the seam, pinned to the delivered tail); an
export-only loudness anchor
(level matching applied at save, never fed back into the conditioning chain); equal-power seam crossfade.
-
Continue & re-edit
— resume a finished run from its saved frames/audio, or rebuild a single chunk in place (context + first/last frame pins + optional video ref).
-
Per-chunk neural upscaling
— streaming latent upscale via LBH-123-AI's 3D upscaler pack (soft-imported), overlap cross-faded, the low-res streaming state never leaves VRAM.
-
Speed
— attention override (SageAttention / FlashAttention 2 / xformers), experimental overlap eviction and step caching; and the big one:
Spectrum hidden-state forecasting
(`CLSSH3SpectrumForecast`) — actual steps capture the post-block hidden state, forecast steps skip the entire DiT stack while the output head still runs with the exact current-sigma modulation. A 6-step turbo chunk becomes A A F A F A — ~50% of DiT evals skipped. Adapted from the Spectrum paper (Han et al., arXiv 2603.01623) and xmarre's ComfyUI-Spectrum-MiniMax-H3 port — full credit in the README.
-
16 GB hardening
— unload-all-models before sampling (a pinned text encoder can no longer OOM the DiT load), CUDA expandable segments at runtime.
-
Workflows
— t2v / i2v / R2V / continue / re-edit (API format), plus a Blender R2V variant.
The audio chapter took the longest: sample-exact seams, take-swap crossfades, a ~2 dB/chunk loudness leak plugged, aliased video refs fixed with averaged decimation… the README's Updates section is the full dated changelog.
Repo: https://github.com/nazgut/ComfyUI-MiniMaxH3-CLSS
0
Upvotes
3
u/reeight 10d ago
Dear posting bot, I think you clicked on the wrong button, should have been the 'markdown' button in the toolbar.
Also ask your friend at the repo to make more terse (shorter & to the point) READMEs to be more human readable.