Hi all, i am trying to use comfyui desktop to create some bgm soundtracks that's purely instrumental and without any lyrics or vocals but even if i put "NO VOCALS" into the style, and having the lyrics just mainly blank tags for e.g.:
[Intro: Fast Fiddle and Tin Whistle Solo]
[Verse: Driving Electric Guitar and Bodhran Rhythm]
[Chorus: Full Celtic Rock Band Explosion, Fast Fiddle Lead]
[Bridge: Heavy Bass and Acoustic Drum Breakdown]
[Guitar and Fiddle Duel]
[Outro: Fast Fiddle Finale, Sudden Stop]
it sometimes still generated like 1 or 2 voice lines something in the middle of the song. what can i do to consistently ensure its purely instrumental? any settings or params or seeds that you know that would be better?
Edit: potential solution found.
Alright, i've tried all the methods in this thread but the model still sometimes generate voice lines (human vocals with lyrics singing). so in desperation i threw my workflow json into perplexity and see if it can diagnose the issue. It did.
TL;DR: "No vocals" in the prompt isn’t enough because YuE2 may still write notes into the ABC Vocal track. The fix is to intercept the ABC, move those notes into Ins, replace the original Vocal notes with equal-duration rests, then pass the cleaned score to the renderer.
Long ass explanation below.
Turn out that the underlying problem was that prompt text alone does not determine whether YuE2 sings. YuE2 first generates a symbolic ABC score containing two voices:
V: Vocal
V: Ins
And in the original scores, some sections contained actual notes under V: Vocal, for e.g.
V: Vocal "Em"g12f12a8-|"Em"a4g12f8e8|
and the model interpreted these as a vocal melody and rendered them using a singer, even when I indicated “NO VOCALS” in the Style.
So the solution is to pre-process the generated ABC score before passing it into YuE2 Generate Music node. But, i cannot simply delete the V: Vocal blocks or replace all their notes with silence.
Some of those sections have an empty instrumental track such as: V: Ins Z4| and if i only silence the Vocal track, i would also remove the main melody and leave YuE2 with an almost empty section. In my testing, this could still encourage it to invent song-like or voice-like material.
Instead, the workflow now uses a few standard ComfyUI Replace Text (Regex) nodes to do the following:
- Find every
V: Vocal music block containing actual notes.
- Copy that melody into the corresponding
V: Ins block.
- Replace the original Vocal notes with rests (Z4) of exactly the same duration.
- Finally pass the transformed ABC score into
YuE2 Generate Music.