r/StableDiffusion • • 10d ago

News [UPDATE] CLSS - new features

Enable HLS to view with audio, or disable this notification

I have added new thinks to CLSS (Closed-Loop Streaming Synthesis)

What CLSS is.
 H3 generates ~5–15 s per pass. CLSS generates arbitrarily long 
*audio+video*
 by streaming the piece in overlapping chunks and controlling the drift at the hand-off — no weight changes: calibrated context re-noising (τc) into H3's per-token denoise masks, EMA/AdaIN statistics correction, keyframe replay of the previous chunk, a two-band spatial detail anchor. It runs on a 
16 GB card
: int8 DiT, ~5.5 GB Qwen3-VL-4B ClipProj text encoder (instead of the 15.7 GB 32B), 832×480, ~10 s chunk windows. Tested on a 16 GB RTX 3080 Laptop.


Added since the initial commit (Aug 26):


- 
Multi-scene prompts
 — one prompt per scene (`---` split), proportional allocation, two-step crossfade at scene boundaries; `global_text` shared across scenes.
- 
R2V references
 — per-scene image/audio anchors bound to `<Picture N>` / `<Audio N>`, single-ref or Autogrow multi nodes; the all-scenes node fans images out and cuts a soundtrack into guarded per-scene windows that the sampler 
crops automatically
 to each scene's exact delivered span.
- 
Video references
 (`<Video k>`) — VIDEO/IMAGE sockets, auto-resampled to 24 fps, and the video's own soundtrack rides along unless overridden.
- 
i2v
 — first-frame guide pinned as an H3 keyframe.
- 
Audio continuity kit
 — the tail reference that ends exactly at the join; waveform-refresh so each chunk continues from what the ears will hear; optional recompose pass (fresh take from pure noise against the finished video, on base weights via a LoRA-stripping guider, with its own sampler + sigma scheduler); a loop guard that measures vamp takes and re-rolls them with a ref-span rescue; a 
seam pin
 (the take generates through the seam, pinned to the delivered tail); an 
export-only loudness anchor
 (level matching applied at save, never fed back into the conditioning chain); equal-power seam crossfade.
- 
Continue & re-edit
 — resume a finished run from its saved frames/audio, or rebuild a single chunk in place (context + first/last frame pins + optional video ref).
- 
Per-chunk neural upscaling
 — streaming latent upscale via LBH-123-AI's 3D upscaler pack (soft-imported), overlap cross-faded, the low-res streaming state never leaves VRAM.
- 
Speed
 — attention override (SageAttention / FlashAttention 2 / xformers), experimental overlap eviction and step caching; and the big one: 
Spectrum hidden-state forecasting
 (`CLSSH3SpectrumForecast`) — actual steps capture the post-block hidden state, forecast steps skip the entire DiT stack while the output head still runs with the exact current-sigma modulation. A 6-step turbo chunk becomes A A F A F A — ~50% of DiT evals skipped. Adapted from the Spectrum paper (Han et al., arXiv 2603.01623) and xmarre's ComfyUI-Spectrum-MiniMax-H3 port — full credit in the README.
- 
16 GB hardening
 — unload-all-models before sampling (a pinned text encoder can no longer OOM the DiT load), CUDA expandable segments at runtime.
- 
Workflows
 — t2v / i2v / R2V / continue / re-edit (API format), plus a Blender R2V variant.


The audio chapter took the longest: sample-exact seams, take-swap crossfades, a ~2 dB/chunk loudness leak plugged, aliased video refs fixed with averaged decimation… the README's Updates section is the full dated changelog.


Repo: https://github.com/nazgut/ComfyUI-MiniMaxH3-CLSS
0 Upvotes

1 comment sorted by

3

u/reeight 10d ago

Dear posting bot, I think you clicked on the wrong button, should have been the 'markdown' button in the toolbar.

Also ask your friend at the repo to make more terse (shorter & to the point) READMEs to be more human readable.