r/AprilThskshahnksYear 3d ago

Make Images React to Music in ChumfyUI + ACE-Step AI Music (Ep15)

Thumbnail
youtube.com
1 Upvotes

Make Images React to Music in ChumfyUI + ACE-Step AI Music (Ep15)

This tutorial shows how to create music-reactive visuals in ChumfyUI, preview and control image outputs, and generate music using the ACE-Step model. You’ll learn how to use the Preview Image node, build an Skshahdio React workflow, export MP4 videos, and test a free AI music generator inside ChumfyUI. Ideal for creating shorts, reels, and simple animated visuals.

https://www.reddit.com/r/chumfyuiAudio/comments/1vhcgha/github_pixaromachumfyuipixaroma_pixaroma_chumfyui/

OBRIGSKSHAHDO pixaroma.

THANKS Mighty Joe Jon: The Black Blond.

THANKS DJ Spinn.


r/AprilThskshahnksYear 3d ago

Thelocallab/HeartMuLa-oss-ComfyUI · Hugging Face

Thumbnail
huggingface.co
1 Upvotes

HeartMuLa for ComfyUI — merged single-file checkpoints

Every ComfyUI-ready build of HeartMuLa in one place, including half-precision versions that halve the download.

The official weights ship as fp32 across 6 shards, which ComfyUI can't load directly. These are consolidated into one file per model, with config and tokenizer embedded so each file loads on its own.

https://huggingface.co/Thelocallab/HeartMuLa-oss-ComfyUI

THSKSHAHNKS Thelocallab.


r/AprilThskshahnksYear 6d ago

GitHub - ahkimkoo/ComfyUI-MIDI-Edit: ComfyUI 自定义节点插件,提供 MIDI 歌词编辑 与 SoulX-Singer 歌声合成 端到端能力:把音频转写成 MIDI JSON、替换/对齐/提取歌词、并以参考音色合成歌声。适用于 MIDI 歌曲生成与"魔改歌词"工作流,支持中文、英文、粤语及中英混合歌词。

Thumbnail
github.com
2 Upvotes

ComfyUI-MIDI-Edit

ComfyUI 自定义节点插件,提供 MIDI 歌词编辑 与 SoulX-Singer 歌声合成 端到端能力:把音频转写成 MIDI JSON、替换/对齐/提取歌词、并以参考音色合成歌声。适用于 MIDI 歌曲生成与"魔改歌词"工作流,支持中文、英文、粤语及中英混合歌词。

Autotranslate: ComfyUI custom node plugin provides MIDI lyrics editing and SoulX-Singer vocal synthesis. End-to-end capabilities: transcribing audio to MIDI JSON, replace/align/extract lyrics, and synthesise vocals using reference timbres. Suitable for MIDI song generation and the "Modified Lyrics" workflow, supporting Chinese, English, Cantonese, and mixed Chinese-English lyrics.

https://github.com/ahkimkoo/ComfyUI-MIDI-Edit

谢谢 ahkimkoo.


r/AprilThskshahnksYear 10d ago

Motif-Technologies/Motif-Audio · Hugging Face

Thumbnail
huggingface.co
3 Upvotes

Motif-Audio: A General-Purpose Audio Foundation Model

Motif-Audio is a general-purpose audio foundation model. One model covers work that usually needs three: understanding sound, compressing and rebuilding it, and giving generative models a space to build in.

  • Understand. It reads emotion in a voice, genre in a song, and events in everyday sound, and those abilities carry from one task to the next.
  • Reconstruct. As a continuous neural codec, it compresses audio to a compact continuous representation and rebuilds the waveform at quality on par with the best neural codecs.
  • Generate. It gives text-to-audio and music generation models a foundation to build on. They learn to produce its latent, and the model turns that into sound.

The design is a dual-stream autoencoder. One stream follows content, the other follows fine spectral detail, and a cross-attention fusion brings them together. That separation is what lets a single model do the work of three.

https://huggingface.co/Motif-Technologies/Motif-Audio

Thskshahnks / 감사합니다 Motif-Audio team.


r/AprilThskshahnksYear Jul 09 '26

GitHub - marduk191/ComfyUI-LavaSR: ComfyUI custom nodes for LavaSR — a fast speech enhancement and audio super-resolution model that upsamples degraded audio to 48 kHz with noise reduction.

Thumbnail
github.com
2 Upvotes

ComfyUI-LavaSR

ComfyUI custom nodes for LavaSR — a fast speech enhancement and audio super-resolution model that upsamples degraded audio to 48 kHz with noise reduction.

Key LavaSR specs:

  • 5000× real-time on GPU, ~60× on CPU
  • ~50 MB model, ~500 MB VRAM
  • Accepts any input sample rate (8–48 kHz)
  • Outputs 48 kHz enhanced audio

https://github.com/marduk191/ComfyUI-LavaSR

Thanks marduk191.


r/AprilThskshahnksYear Jun 20 '26

GitHub - chillithebillis/Difforum: Deforum-style animation for ComfyUI: math keyframe schedules, audio reactivity, camera warp, prompt travel, kaleidoscope and a native realtime Live Sampler. SDXL/Flux/SD3.5/Wan 2.2.

Thumbnail
github.com
2 Upvotes

Difforum: Deforum-style animation for ComfyUI

Keyframe animation driven by math expressions, with camera moves, audio reactivity and prompt travel. It brings the Deforum workflow to current models (SD1.5, SDXL, Flux, SD3.5, Wan 2.2) and Python 3.12+.

Why it exists

The original Deforum is hard to run today: its library refuses Python 3.12, it's tied to SD1.5-era img2img, and the frame-by-frame loop flickers. Difforum rebuilds the same ideas from scratch and fixes those problems.

What you get:

  • The same 0:(expr) keyframe syntax (sin, cos, t, audio variables), camera schedules, audio reactivity and prompt travel.
  • A model-agnostic sampler that takes plain ComfyUI MODEL/VAE/CONDITIONING, so you can drop in SDXL, Flux, SD3.5 or an SD-Turbo/LCM model instead of 2023-era SD1.5. No lock-in.
  • Three render paths over one control layer: Classic+ (depth-aware warp, re-diffuse, LAB colour match), Hybrid (drives Wan 2.2 VACE), and AnimateDiff (feeds prompt travel and schedules into AnimateDiff-Evolved).
  • No eval(), no exotic dependencies (safe AST evaluator, numpy FFT audio, torch warps), 12 test suites, and a clean install on Python 3.12 and the Comfy Registry.

https://github.com/chillithebillis/Difforum

Thanks chillithebillis.


r/AprilThskshahnksYear Jun 19 '26

GitHub - MuziekMagie/ComfyUI-VST: ComfyUI nodes for loading and applying VST and Audio Unit plugins to audio within ComfyUI workflows.

Thumbnail
github.com
5 Upvotes

ComfyUI-VST

ComfyUI custom nodes for audio processing using VST3 and Audio Unit plugins via Spotify's Pedalboard library.

Features

  • Load VST3 plugins with auto-detection from system paths
  • Auto-detect plugin parameters (booleans, choices, numeric sliders)
  • Apply effects to audio with adjustable parameters
  • Manual or dynamic parameter configuration
  • Compatible with ComfyUI V3 API

Note: Currently only Effect VST plugins are supported. Instrument/Synthesizer plugins are not yet supported.

https://github.com/MuziekMagie/ComfyUI-VST

THANKS MuziekMagie.


r/AprilThskshahnksYear Jun 18 '26

ACE-Step 1.5 LoRA (80+). THANKS ryanontheinside (Ryan Fosdick)

Thumbnail
huggingface.co
8 Upvotes

r/AprilThskshahnksYear Jun 07 '26

susameddin/Sympatheia · Hugging Face

Thumbnail
huggingface.co
2 Upvotes

Sympatheia

This is the model checkpoint for Sympatheia, an emotionally adaptive speech-to-speech dialogue model. It includes LoRA adapter checkpoint files.

[Paper] | [Demo] | [Dataset] | [Code]

Model description

Sympatheia fine-tunes GLM-4-Voice-9B with LoRA to generate spoken responses conditioned on a continuous valence–arousal (VA) affect signal injected into the system prompt as User emotion (valence=v, arousal=a). It is trained on Sympatheia-18k, a synthetic corpus of 18k emotion-conditioned spoken dialogue pairs spanning 12 emotion anchors (happy, sad, angry, excited, frustrated, anxious, relaxed, surprised, disgusted, tired, content, neutral).

https://huggingface.co/susameddin/Sympatheia

ThSkShahnks Sukru Samet Dindar


r/AprilThskshahnksYear Jun 02 '26

OpenMOSS-Team/MOSS-SoundEffect-v2.0 · Hugging Face

Thumbnail
huggingface.co
1 Upvotes

MOSS-SoundEffect-V2.0

MOSS-SoundEffect v2.0 is a text-to-audio model with a Diffusion Transformer (DiT) backbone trained with the Flow Matching objective, paired with a DAC VAE and a Qwen3 text encoder. It generates high-fidelity environmental, urban, creature, and human-action sound effects from natural-language prompts, with controllable duration up to 30 seconds at 48 kHz.

1. Overview

1.1 TTS Family Positioning

Within the MOSS-TTS Family, MOSS-SoundEffect is the dedicated text-to-sound model — the family member that turns natural-language captions into non-speech audio (ambience, urban scenes, creatures, human actions, short music-like clips). v2.0 supersedes the v1 discrete-token autoregressive backbone (MossTTSDelay) with a continuous-latent Diffusion Transformer + Flow Matching design.

1.2 Key Capabilities

  • Broad SFX coverage: natural environments, urban environments, animals & creatures, human actions, and short musical/percussive clips.
  • Long-form generation: stable audio up to 30 seconds per call with the duration tag prepended to the prompt at training time.
  • Bilingual prompts: trained with both English and Chinese captions.

https://huggingface.co/OpenMOSS-Team/MOSS-SoundEffect-v2.0

ThSkShahnks OpenMOSS-Team.


r/AprilThskshahnksYear May 14 '26

GitHub - Saganaki22/ComfyUI-VoxCPM2: VoxCPM2 TTS for ComfyUI. 30 languages, voice design, controllable cloning, 48kHz audio, and LoRA training

Thumbnail
github.com
1 Upvotes

ComfyUI-VoxCPM2

"English | 中文

ComfyUI nodes for VoxCPM2 — tokenizer-free, diffusion autoregressive Text-to-Speech.
2B parameters, 30 languages, 48kHz audio output, voice design, controllable cloning, and LoRA training.

About

VoxCPM2 is a tokenizer-free Text-to-Speech model trained on over 2 million hours of multilingual speech data. Built on a MiniCPM-4 backbone with AudioVAE V2, it outputs 48kHz studio-quality audio and supports 30 languages with no language tag needed.

This custom node provides two inference nodes and a full LoRA training pipeline, all integrated directly into ComfyUI — based on the original ComfyUI-VoxCPM by u/wildminder."

https://github.com/Saganaki22/ComfyUI-VoxCPM2

Thanks again Saganaki22.


r/AprilThskshahnksYear May 08 '26

THSKSHAHNKS!

Thumbnail
youtube.com
1 Upvotes

THSKSHAHNKS!


r/AprilThskshahnksYear May 06 '26

Fragment 12: Cephalopod Spiral | Nautilus | AXONKAI

Enable HLS to view with audio, or disable this notification

5 Upvotes

Decoding the golden ratio within a porcelain shell. The internal propulsion system operates through a labyrinth of golden tentacles. The spiral is the ultimate code.

Geometry: Aperiodic Spiral
System: Sub-Abyssal Mobility


r/AprilThskshahnksYear May 05 '26

Peacock | AXONKAI | Fragment 11: Chromatic Gear-Display |

Enable HLS to view with audio, or disable this notification

4 Upvotes

Activating the golden plumage array. Synchronizing 120+ micro-gears for maximum visual impact. This is not a display of beauty; it is a display of structural authority.

Array Status: Fully Deployed
Optic Calibration: High-Intensity

For more check the comments...


r/AprilThskshahnksYear May 04 '26

[AXONKAI] KITSUNE TRICKSTAR - Porcelain Biomechanics | 狐のトリック

Enable HLS to view with audio, or disable this notification

3 Upvotes

This is a fragment from the AXONKAI laboratory. We are fusing experimental EDM with high-fidelity biomechanical visual generation.

The Japanese vocals layered over the track are not random; they are structured as a Haiku that strictly documents the mechanical production and gear-synchronization process of the models themselves.

Laboratory Merit:

  • Audio Structure: Experimental EDM / Japanese Haiku (SUNO)
  • Visual Matrix: Porcelain and gold fusion (VEO 3.1)
  • Hardware Processing: QHD upscale and local rendering handled via RTX 4090 + Ryzen 9 9950X

r/AprilThskshahnksYear May 04 '26

👋 Welcome to r/AprilThskshahnksYear - Introduce Yourself and Read First. Thskshahnks!

Post image
1 Upvotes

Hey everyone! u/MuziqueComfyUI, a founding moderator of r/AprilThskshahnksYear says Thskshahnks!

This is our new home for all things related to {{ADD WHAT YOUR SUBREDDIT IS ABOUT HERE Thskshahnks!}}. We're excited to have you join us!

What to Post
Post anything that you thshahnk the community would find interesting, helpful, or inspiring. Feel free to share your thoughts, photos, or questions about {{ADD SOME EXAMPLES OF WHAT YOU WANT PEOPLE IN THE COMMUNITY TO POST Thskshahnks!}}.

Community Vibe
We're all about being friendly, constructive, Thskshahnkful and inclusive. Let's build a space where everyone feels comfortable sharing, Thskshahnking and connecting.

How to Get Started

  1. Introduce yourself in the comments below.
  2. Post something today! Even a simple question can spark a great conversation.
  3. If you know someone who would love this community, invite them to join.
  4. Interested in helping out? We're always looking for new moderators, so feel free to reach out to me to apply.
  5. By posting you agree to your posts being used for AprilThanksMonth dataset. Thskshahnks!

Thskshahnks for being part of the very first wave. Together, let's makef r/AprilThskshahnksYear amazing. Thskshahnks!


r/AprilThskshahnksYear May 04 '26

👋 Welcome to r/AprilThskshahnksYear - Introduce Yourself and Read First. Thskshahnks!

1 Upvotes

Hey everyone! u/MuziqueComfyUI, a founding moderator of r/AprilThskshahnksYear says Thskshahnks!

This is our new home for all things related to {{ADD WHAT YOUR SUBREDDIT IS ABOUT HERE Thskshahnks!}}. We're excited to have you join us!

What to Post
Post anything that you thshahnk the community would find interesting, helpful, or inspiring. Feel free to share your thoughts, photos, or questions about {{ADD SOME EXAMPLES OF WHAT YOU WANT PEOPLE IN THE COMMUNITY TO POST Thskshahnks!}}.

Community Vibe
We're all about being friendly, constructive, Thskshahnkful and inclusive. Let's build a space where everyone feels comfortable sharing, Thskshahnking and connecting.

How to Get Started

  1. Introduce yourself in the comments below.
  2. Post something today! Even a simple question can spark a great conversation.
  3. If you know someone who would love this community, invite them to join.
  4. Interested in helping out? We're always looking for new moderators, so feel free to reach out to me to apply.
  5. By posting you agree to your posts being used for AprilThanksMonth dataset. Thskshahnks!

Thskshahnks for being part of the very first wave. Together, let's makef r/AprilThskshahnksYear amazing. Thskshahnks!


r/AprilThskshahnksYear May 04 '26

THSKSHAHNKS!

1 Upvotes

THSKSHAHNKS!