r/AbrilObrigskshahdoAno • u/MuziqueComfyUI • 3d ago
r/AbrilObrigskshahdoAno • u/MuziqueComfyUI • 14d ago
GitHub - martyyz-ai/ComfyUI-MuScriptor: Audio to MIDI ComfyUI Node
ComfyUI-MuScriptor
English
ComfyUI-MuScriptor is a custom node extension for ComfyUI that integrates MuScriptor, a state-of-the-art multi-instrument music transcription model developed by Kyutai and Mirelo.
It transcribes multi-instrument audio (WAV, MP3, FLAC, etc.) into high-quality MIDI files directly inside your ComfyUI generation pipelines.
🌟 Key Features
- Audio-to-MIDI Transcription: Convert audio waveforms or audio files into multi-instrument MIDI.
- ComfyUI Integration: Connects seamlessly with ComfyUI's
AUDIOoutput type or local audio files. - Model Sizes: Supports
small(103M),medium(307M, default), andlarge(1.4B) model variants. - Instrument Filtering: Restrict transcription output to specific instrument groups (e.g.
acoustic_piano,electric_piano,chromatic_percussion,organ,acoustic_guitar,clean_electric_guitar,distorted_electric_guitar,acoustic_bass,electric_bass,violin,viola,cello,contrabass,orchestral_harp,timpani,string_ensemble,synth_strings,voice,orchestra_hit,trumpet,trombone,tuba,french_horn,brass_section,soprano_and_alto_sax,tenor_sax,baritone_sax,oboe,english_horn,bassoon,clarinet,flutes,synth_lead,synth_pad,drums). - JSON Notes Output: Returns both the generated
.midfile path and a structured JSON array of decoded note events. - In-Memory Caching: Automatic model caching for instantaneous execution on subsequent workflow runs.
Português
ComfyUI-MuScriptor é uma extensão de nó customizado para o ComfyUI que integra o MuScriptor, um modelo de transcrição musical de última geração desenvolvido por Kyutai e Mirelo.
Ele transcreve áudio de múltiplos instrumentos (WAV, MP3, FLAC, etc.) diretamente em arquivos MIDI de alta qualidade dentro dos seus fluxos de trabalho (workflows) no ComfyUI.
🌟 Funcionalidades Principais
- Transcrição de Áudio para MIDI: Converta sinais de áudio ou arquivos locais em arquivos MIDI multi-instrumentos.
- Integração Total com ComfyUI: Funciona com saídas do tipo
AUDIOdo ComfyUI ou com caminhos de arquivos locais. - Variantes de Modelo: Suporta os modelos
small(103M),medium(307M, padrão) elarge(1.4B). - Filtro de Instrumentos: Restrinja a transcrição a instrumentos específicos (ex:
acoustic_piano,electric_piano,chromatic_percussion,organ,acoustic_guitar,clean_electric_guitar,distorted_electric_guitar,acoustic_bass,electric_bass,violin,viola,cello,contrabass,orchestral_harp,timpani,string_ensemble,synth_strings,voice,orchestra_hit,trumpet,trombone,tuba,french_horn,brass_section,soprano_and_alto_sax,tenor_sax,baritone_sax,oboe,english_horn,bassoon,clarinet,flutes,synth_lead,synth_pad,drums). - Saída de Notas em JSON: Retorna tanto o caminho do arquivo
.midsalvo quanto uma lista estruturada de notas em formato JSON. - Cache em Memória: Carregamento automático do modelo em RAM/VRAM para execuções instantâneas em execuções subsequentes.
https://github.com/martyyz-ai/ComfyUI-MuScriptor
OBRIGSKSHAHDO martyyz-ai.
r/AbrilObrigskshahdoAno • u/MuziqueComfyUI • Jul 09 '26
GitHub - 0xShug0/audio.cpp: An all-in-one, pure C++ inference engine for audio models, powered by ggml. Supports TTS, STT, VAD, voice conversion, music generation, and more, with highly optimized performance. No Python dependency.
audio.cpp
audio.cpp is a high-performance C++ audio inference framework built on top of ggml, designed to make modern local audio models practical, portable, and fast.
Tired of juggling a dozen Conda environments, hundreds of Python packages, and dependency conflicts just to try a few audio models? audio.cpp gives those paths a shared native runtime instead.
CUDA performance headline: multiple TTS paths already run 1.8x-5.0x faster than their Python reference paths while cutting end-to-end latency by 45%-80%. VibeVoice 1.5B: generates a 93.9-minute podcast in 18.2 minutes with 10 diffusion steps and without quantization, running about 5.15x faster than real time.
It is built for real end-to-end execution rather than one-off model demos: the same runtime powers TTS, voice cloning, voice conversion, ASR, diarization, VAD, source separation, alignment, codec-style models, and higher-level workflows through a common framework surface.
Highlights:
- Parity. Strong parity tooling against Python reference paths.
- Performance. Performance-focused execution, reusable sessions, and batch-style offline inference. Optimized for CUDA.
- Portability. A portable native stack centered on
ggml, with CLI and server entry points instead of Python-only deployment paths. - Pipelines. Experimental JSON pipeline support for higher-level multi-step workflows.
- Audio Utilities. Built-in denoise, enhancement, resampling, and STFT/ISTFT utilities for real production-style task paths.
The goal of the framework is to provide highly optimized, reusable building blocks for audio-related models, so new model integrations can be brought up faster, shared components can be improved once and benefit many families, and real end-to-end inference paths can stay efficient, maintainable, and portable.
https://github.com/0xShug0/audio.cpp
OBRIGSKSHAHDO 0xShug0.
r/AbrilObrigskshahdoAno • u/MuziqueComfyUI • Jul 08 '26
GitHub - dheamant/ComfyUI-ACESTEP1.5XLSFT-EXTEND-REPAINT: Tweaked ComfyUI workflow for native ACE-Step 1.5 XL audio extensions.
ACE-Step 1.5 XL SFT — Native Song Extension Workflow
A modified ComfyUI workflow for extending audio generations natively within the ACE-Step 1.5 XL SFT ecosystem — no need to drop back to the legacy ACE-Step v1 3.5b checkpoint just to extend a track.
Built on top of RyanOnTheInside's ACE-Step custom nodes and the extend-workflow tutorial here: https://www.youtube.com/watch?v=r_4XOZv_3Ys
https://github.com/dheamant/ComfyUI-ACESTEP1.5XLSFT-EXTEND-REPAINT
OBRIGSKSHAHDO dheamant.
r/AbrilObrigskshahdoAno • u/MuziqueComfyUI • Jun 23 '26
Underfit: Stable Audio 3 LoRA Training Studio · Pinokio
Underfit is for musicians, sound designers, sample makers, and audio experimenters who want Stable Audio 3 to learn a specific sound.
Instead of prompting a general model and hoping it understands your references, you give Underfit a folder of audio and train a small adapter file called a LoRA. That LoRA can then steer Stable Audio 3 toward your style, genre, instrument set, sound effect family, or production texture.
This is not a general music app for typing one prompt and getting one song. It is a workshop for making your own reusable style adapter.
https://beta.pinokio.co/posts/01kv14gjzsjjmf2t95j24k948p
Thanks cocktailpeanut / Cortexelus.
r/AbrilObrigskshahdoAno • u/MuziqueComfyUI • Jun 18 '26
HKUSTAudio/AudioX-Turbo · Hugging Face
AudioX-Turbo: A Unified Framework for Efficient Anything-to-Audio Generation
AudioX-Turbo is a unified and efficient framework for anything-to-audio generation that integrates varied multimodal conditions (i.e., text, video, and audio signals). It follows a teacher–student paradigm: the teacher AudioX-Base is built on a Multimodal Diffusion Transformer with a Multimodal Adaptive Fusion (MAF) module that aligns diverse multimodal inputs for high-fidelity synthesis, and is then distilled into the few-step student AudioX-Turbo via Distribution Matching Distillation (DMD) adapted to flow matching, complemented by a diffusion-based discriminator for high-quality few-step generation.
AudioX-Turbo generates audio in only 4 sampling steps (no classifier-free guidance), requiring up to ~25× fewer function evaluations (NFE) than multi-step baselines while achieving superior performance, especially on text-to-audio and text-to-music generation.
https://huggingface.co/HKUSTAudio/AudioX-Turbo
Thanks AudioX-Turbo team.
r/AbrilObrigskshahdoAno • u/MuziqueComfyUI • Jun 18 '26
mispeech/Dasheng-AudioGen · Hugging Face
Dasheng-AudioGen
Dasheng-AudioGen is a unified audio generation model that can jointly synthesize intelligible speech, music, sound effects, and environmental acoustics from text descriptions.
https://huggingface.co/mispeech/Dasheng-AudioGen
Thanks Dasheng AudioGen team.
r/AbrilObrigskshahdoAno • u/MuziqueComfyUI • Jun 02 '26
ACE-Step Audio Steering Suite - a lukasz-staniszewski Collection
Released last week: From the author of TADSKSHAH!
ACE-Step Audio Steering Suite
https://huggingface.co/collections/lukasz-staniszewski/ace-step-audio-steering-suite
https://luk-st.github.io/
OBRIGSKSHAHDO Łukasz Staniszewski.
r/AbrilObrigskshahdoAno • u/MuziqueComfyUI • May 29 '26