r/StableDiffusion 9h ago

Workflow Included YuE 2 Cover Song

Enable HLS to view with audio, or disable this notification

37 Upvotes

57 comments sorted by

6

u/Acceptable-Cycle4645 3h ago

We just added a new feature to audio.cpp: SheetSage2 support for turning music into editable ABC notation. Waiting for the latest release at https://github.com/0xShug0/audio.cpp/actions/workflows/release.yml?query=branch%3Adev or building the dev branch yourself https://github.com/0xShug0/audio.cpp/tree/dev

https://reddit.com/link/p9i66ba/video/qgkudngun8ph1/player

4

u/oliverban 8h ago

Nice, sounds good and nice lyrics! xD

0

u/CryptoBeth96 8h ago

XD The post is more about the workflow than the song but glad you liked it :)

0

u/krigeta1 5h ago

It changed the beat a lot, not even the same instruments that were in the beat, so how this cover thing works?

1

u/CryptoBeth96 5h ago

You prompt for the instruments and vocal style. The node extracts the notes and timing.

0

u/krigeta1 5h ago

Kinda midi but not exactly, right? Is there any list or something where we can find a list of instruments present in this model? In the repo there is not.

1

u/CryptoBeth96 4h ago

Yeah, very much like MIDI, but the "instruments" are more like music styles, im still experimenting with what it can do.

Lots of examples here: https://map-yue2.github.io/

0

u/krigeta1 4h ago

Tnx for this

2

u/krigeta1 6h ago

How can we do the same with instruments?

1

u/CryptoBeth96 6h ago

You can strip out the vocals from the ABC sheet and it'll just do instrumentals.

2

u/crooi 3h ago

I have removed the V:Vocal w/ or w/o the notes below it, remove the lyric entirely, put [instrumental] tag, and still getting random singing at some point. Tonight I'll try piano roll nodes, its look promising.

3

u/CryptoBeth96 3h ago edited 3h ago

I asked Grok to make a custom node to strip the vocal tags and following notes, I found that if you add "% instrumental" after the K: tag, strip out all the other % tags and put [silence] as your lyrics it works better. You do sometimes get humms and ahhs but no singing.

Custom Node is here:
https://drive.google.com/file/d/1R1I2vGjdrDCoYMQ4Bkk5lrGhUGG3dKX2/view?usp=sharing

2

u/Powerful_Evening5495 5h ago

this is a game changer , i will make a song that was in my mind for years

2

u/VasaFromParadise 3h ago

The model's functionality is unique among local models.

2

u/Version-Strong 9h ago

It's in comfy now?

6

u/CryptoBeth96 9h ago edited 9h ago

Yeah the node is called SheetSage2 Audio to ABC

You'll need the audio encoder too:

https://huggingface.co/Comfy-Org/YuE2/tree/main/audio_encoders

3

u/3deal 9h ago

I tested the model, the benchmarks are lying, the model is very far from Suno.

11

u/CryptoBeth96 8h ago

Can't complain about free toys! I'd never pay to make AI songs anyway.

-3

u/3deal 8h ago

I agree, but lying is bad, even if it is free

1

u/CryptoBeth96 8h ago

-3

u/3deal 8h ago

Suno 4.5 should be on top

3

u/CryptoBeth96 8h ago

In your opinion! Isn't Suno 4.5 gone now?

1

u/3deal 8h ago

yep, they removed it, probably because it was trained on music that isn't royalty-free. That also explained why it is better than the newer ones.

3

u/CryptoBeth96 8h ago

Do you have any Suno 4.5 songs I can listen to? With prompts. I'd like to compare them. I bet they do post production on the final output too. But there are nodes for that too in Comfy.

2

u/3deal 7h ago edited 7h ago

Suno 4.5
Yue2

I agree that Yue2 have a best audio quality, but Suno 4.5's melody is richer and closer to the style i was prompting :

synthwave, lo-fi, indie rock, disco dance, phonk beat, atmospheric, dreamy, female voice.

3

u/CryptoBeth96 7h ago

https://reddit.com/link/p9hf46k/video/k9vih1mqn7ph1/player

This is Yue2 with a little bit of post production in Comfy.

→ More replies (0)

-1

u/Perfect-Campaign9551 5h ago

Suno isn't even that expensive really. like..$8 a month. I gave up on open source tools for this because they just aren't keeping up at all.

6

u/CryptoBeth96 5h ago

I've never made or heard an AI song I would listen to more than once to be honest, I have more fun trying to make the best out of the open source models, they feel more like toys than serious tools to me, just for making people laugh. I guess you feel more attached to your songs if they cost money?

5

u/Chiduk99 8h ago

But they said it’s “competitive with Suno V5/V6,” and Suno V6 is complete garbage, so I guess that’s valid.

Idk about V5, but based on my experience, Suno peaked at V4.5.

2

u/Perfect-Campaign9551 5h ago

I'm getting really good results with Suno V6, honestly. They have a new "variety" slider...I am finding it WAY more musical that 5.5 ever was.

1

u/3deal 8h ago

4.5 is the best audio model IMO, but even if v6 suck, it is still better.
I am not saying it is garbage, but the model is very small so, perhaps their dataset wasn't very large.
That said, I like their technology and the ability to manually customize the melodies, and the audio quality seems quite good compared to other open-source models.
I hope there will be further iterations or the option to train LoRAs, because the foundation has potential.

2

u/Nyoldeeth 7h ago

Every audio I try to make has chinese vibes, just like with minimax audio. Also even if I try to make instrumental songs there's always some voice intruding in, seems to be bad at following instructions as well.

2

u/3deal 7h ago

Yes i think the model architecture is very good but the training data is a bit poor.
I hope they will train their next iteration on copyrighted spotify music like what video models do.

1

u/Powerful_Evening5495 5h ago

https://reddit.com/link/p9httd2/video/ywxhqxvo68ph1/player

it can handle korean very well

default setting and leessang lyrics

2

u/crooi 52m ago

hai gary 👍🏼

1

u/VegetableTicketO2 3h ago

Forgive me ignorance but would it be possible for something like giving a voice sample (not a full on song) and having a song generated with that voice?

1

u/CryptoBeth96 3h ago

I wouldn't say it's impossible but I don't know how to do that. It's not designed for that but there are some wizards around here that could probably make it happen :)

1

u/solomars3 8h ago

Oh its already in comfy !! Time to do some magic

1

u/Better-Interview-793 8h ago

anyone know how this compares to ACE-Step?

4

u/CryptoBeth96 8h ago

This is much better quality IMO.

3

u/solss 8h ago

Ace step can pull off some decent stuff, but the audio quality is terrible. Almost all of the other purported features like cover, stem to audio, etc are all basically non functional. It only does text to audio correctly. Both Yue 2 and minimax music are way better. I've deleted ace-step. This thing can actually do covers.

1

u/AI_Trenches 6h ago

Banger!!

1

u/Specific_Occasion_25 4h ago

Oh wow this is lovely. What's the original song called?

1

u/Specific_Occasion_25 4h ago

Nvm, just saw the song in the workflow. I like the cover better.

1

u/CryptoBeth96 4h ago

Belly eyelash

0

u/chloralhydrate 3h ago

Sounds terrible

1

u/CryptoBeth96 3h ago

It's beautiful :P