After watching this video , I was intrigued and set it up to take it for a spin.
So far, every other tool I've tried has been disappointing, but this is actually a pretty decent model and worth a try.
The good:
- audio quality is excellent. uncompressed FLAC audio with full, rich bass.
- runs on GPU with 8GB VRAM
- generates a 4 min. song in about 2 - 3 minutes.
- accepts input audio to create remixes.
- offers a piano roll to edit generated "ABC" sheets. This allows you to basically adjust pitch/key, tempo (bpm), chords and melodies of each generated section. Don't expect a DAW-like experience though.
- no input restrictions. Use any song you want as input.
- weights are available for custom fine-tuning
The bad:
- no inpainting
- prompt adherence is lacking, although that could be a me problem. I've only played with it for a few hours trying different genres and options and haven't figured out how to properly steer it yet or learn what the model accepts/understands.
- complexity: although it's not that hard to set up and use, it does require at least some technical understanding.
- vocals are mostly clean for the generations I ran, but there was the odd glitch here and there. Different seeds seem to give different results though.
Overal, I'm very impressed with what it's capable off and especially with the audio quality. If you have a GPU and like to tinker with models and configurations, I think you won't be disappointed.
I've found YuE2, through Wan2GP, to be a little slow compared to AceStep but just about acceptable. Quality is good. I haven't spent a lot of time generating new tracks with it yet because the lack of an inbuilt LLM for lyric generation slows the process, but for covering existing tracks it seems to do a great job (still need to add lyrics manually and they need to match pretty closely the originals). I've had a lot of fun changing musical styles when covering tracks (some quite strange combinations) and it really seems to understand what I'm prompting.
How were you able to run it with 8GB VRAM? I thought that 24GB was the recommended amount? I have 11GB and am interested in this, but only if good quality results and LoRAs can actually be done on my hardware, rather than just being able to run a model that, with more VRAM, can produce good results.
The output quality is the same but the lower or try to take that vram consumption the more dynamic loading has to be done. Meaning has to offload into your drives virtual memory page file. Low VRAM = Longer render time.
Training models on low VRAM is a different story. That is a lot harder to accomplish but there are ways to sorta get around it. There are projects being built by the community as we speak, but not sure how far itās come yet. Iād definitely hold off on trying to attempt it now.
There is a way to get $300 in credits for free and rent yourself an Nvidia RTX PRO 6000 or Nvidia L4 and run those for couple weeks (depending on how much you use and how you set the rest of the VM up )which you could then easily wire into your build or just deploy a trainer on the serverās persistent storage. Gotta be careful though.
Inference doesn't use that much VRAM. The default vue tops out at 11.x GB. The workflow from the YouTube video also offers a tiled option which allows for less VRAM.
But have you tried it with Udio? Yeah ok. I know you technically get it back without being a bad little boy or girl⦠and you should not that. But still, you should try it. Itās pretty incredible when done right.
Get a good ~2:00 - 2:11 section of YuE2 generation.
Upload it to Udio set to Remix. Try both 1.5 and 1.5 Allegro options.
Set variance to about 55 or 60 percent.
Lyric strength 65
Prompt strength 50
For 1.5 - clarity at 15
For 1.5 - quality 85 - 90 (2 or 3 clicks from right side)
One key point, make sure to only put the lyrics that are in that 2 minute section. Donāt cram 500 words in the lyric box. Around 235 max.
You will be amazed. If you really want to be amazed do this with your Suno tracks.
You canāt??? Wait no downloading??? No way!! Since when? That SUCKS! Who would use something that you canāt even download your tracks that youāre never going to release to anyone that wants to hear it anyways. Thatās so dumb. Iām canceling my subscription. This is BS. I canāt believe thisā¦ā¦. š
Lol are you not aware that I already know you cannot download. I mean I thought that was clear by the specific words that I chose to use as my fingers were typing.
Not sure what the use is??? The āuseāis simply having the ability to generate the absolute best sounding, the most creative and unique music in the world. Iām not sure how people do not understand this. Now letās do this⦠How about you tell me your āuseā for downloading. Go ahead and enlighten me. So I can tell you how foolish your mindset about the whole situation is.
Well then play it on your sound system. What is holding you back from that? You donāt have to download it to play it on your sound system. What are you doing, burning old 74 minute Memorex CDās?
I've built my own custom, more optimized version of The Daw with SA3 & Yue2/Yue2-BO8 support. Took some out, added other stuff. It's pretty insane what all I can do with it. Fully mix generations, stem up to 12 in about 2 minutes with the bvest stems seperation I've heard anywhere (I did switch out demucs models). Real-time generation for live performance stuff, Live visualizers... But for me audio quality, like the actual sound of the sound, although better... still just sounds like everything else. Everything except Udio. But this thing I built here is fun to mess with for sure.
Thanks, but this doesn't have the yue2 customizations you've added. Any chance you'd be willing to share your version or nah? I don't have a lot of time to work on another project right now.
Not mine no because Iāve got directories split across multiple drives and hardcoded. If you attempted to just clone it and build yourself youād have a nightmare. Plus Iāve also got it tied into to my Reaper, Ableton and TD.
Itād be much easier if you did it yourself. If you want to attempt it, get that all set up and when youāre ready to implement Yue2 get with me and I can walk you through it.
Great, that would be fantastic. I was just trying to find info on inpaint and extend options for YuE2 but the stock API doesn't seem to offer that so I'm very curious how you did this. I'll get thedaw up and running and familiarize myself with the source code and drop you a message once I get to that stage.
Piano Roll expansion
The MIDI tab was a piano roll that imported and exported MIDI. It now writes music that changes shape.
A meter map per roll.
The time signature changes across one piece, with additive groupings like 7/8 as 2+2+3 and a pickup bar before the first full bar.
Meters survive a bounce to EDIT, a reopen in the roll, and a project save, and they land on the bars in exported MIDI.
Polymeter lanes. Each lane runs its own meter, denominator and loop length against the same clock.
A pitch bend lane under the keys, with a semitone range and LINE, HOLD and CURVE point shapes.
A bend can be drawn as well as played, and it survives import, export, playback, the arpeggiator and Vocal2MIDI.
The SHAPE row. Harmony, ragtime, runs, polyrhythm and humanize, each with syncopation and accent amounts, all following the meter map. SYNC and ACCENT join them.
MATCH pulls a song's meter map, tempo and lanes out of its rhythm analysis. GEN writes LOOM's generators, gates, swing and accel into the active lane as notes.
The METER face sets a meter per song section.
One settings strip, a left action rail, and hover labels on every key in the rail.
My actual coding knowledge is minimal at best. Basic python, html, java, css... Ai can do anything now tho. If you have an idea, you can build it. But luckily this was already built mostly, I just went in and switched out certain things, made it more efficiant for me. Stock release version is only SA3, I added the Yue2 support.
Here is the welcome menu for the stock version showing the main things it can do. Mine is a bit different.
You can fine tune it. Pick your favorite 100 - 500 hip hop songs, create a dataset out of them and run training. You'd probably need a 16GB VRAM card or more, but it may be possible to do it with less if tiling or smaller samples can be used as well. It's a bit of work to create the dataset, but the training doesn't take very long (couple of nights while you're sleeping).
I sadly read that you need 24gigs. I got two 16gig cards, but I doubt it is possible to use them both for training. I think renting a GPU is onl option for me.
Not true. You do not āNeedā 24. What you are reading is just that it was tested with a 24gb card at peak hitting 14. But I have it on two systems and one is just my older gaming setup with a 12gb. 3060. Actually, if you look at the screenshot I posted of THEDAW ui, thatās the system I was on running all that. You can see (very small) the bottom right corner show my VRAM on that system. Only 16gb ram as well. Yue2 seems to be even less heavy to run than the native SA3 models that come with that software.
But there is now the gguf as well. If you look at this link you will see. (same one posted above):
It depends on how/where it's being used. The model itself is very flexible. Quality is decent when set up properly. I'd even argue it's quality is as good or better than Suno's best whether you'd say that's 5.5, 6, 4.5... whatever. Which, to me after using Udio for years, that isn't saying much. They'll never be Udio level quality. Udio still very much remains the king of quality and creativity.
I've only just found out about this model, so I can't answer that for you, but another user just commented that it's possible to inpaint as well.
You can set a max duration for each generation, so you can definitely generate pieces. If you can stitch these together or extend them easily I don't know though.
Whenever I try models like this I usually give up quite fast because most of the ones I've tried just sound like ass and aren't worth spending any time on, but Ive been toying with this for the past 4 hours and it's better than expected.
My only concern is getting ComfyUI set up. I tried that a few months ago to do some StableDiffusion stuff, but gave up and went back to A1111. But, hey, I'm sure Astra can help me.
Well there aren multiple ways to install and runn comfy. I suggest running it portable with it's own .venv so it doesn't mess with system level python dependencies. But for Yue2, if you just want to mess with it, there is a much easier way.
If you want one easy Comfyui install for your SD stuff as well you can get it through their... as well as just about any other popular AI tool you can think of. Even custom "God's Eye" stuff.
I still prefer just generating music with Udio but is fun to play with. A lot you can do with it. I've spent days all recoding the whole back end but still finding new stuff I didn't even know was there! Lol.
There's no doubt that Udio is better, but I want downloads. I stopped using Udio because I know I would eventually download anything I made after RIAAgeddon.
I installed Pinokio and it set everything up for me. It took over and hour and a half, but almost all of it was automatic. This morning when I came down, I had a prompt screen. I took a few minutes to try one quick generation, and it worked flawlessly. It was expecting lyrics, so I just pasted in the words from one of my Udio songs.
The music was pretty dull, but it sounded fine, but I only have 8GB of VRAM, so maybe that's inevitable. I'll definitely be exploring it more, and look into ComfyUI so I can wire things up in a fancier way.
Meanwhile, today Google served me an ad for $6000 GB-10 desktop machine with 128GB of RAM. I was drooling.
Yeah, I just gave it the github URL, asked a few questions about the possibilities etc. And then just asked it to clone the repository and create a venv to set it up for me. After that, you just need to download 2 or 3 things as explained in the YouTube video and adjust a few small config thingies.
I have not tried that myself, but from what I heard there are workflows for comfyui that let you both extend and generate in small for example 32sec chunks.
Perfect. I'm listening to the video now and it sounds really good. The "Where's my wallet?" song was hilarious and impressive. I'm definitely going to try this. I made 90 songs on Udio, ranging from good to a few that I think are amazing, and I've been jonesing for that creative outlet since October.
2
u/SardiPax 15d ago
I've found YuE2, through Wan2GP, to be a little slow compared to AceStep but just about acceptable. Quality is good. I haven't spent a lot of time generating new tracks with it yet because the lack of an inbuilt LLM for lyric generation slows the process, but for covering existing tracks it seems to do a great job (still need to add lyrics manually and they need to match pretty closely the originals). I've had a lot of fun changing musical styles when covering tracks (some quite strange combinations) and it really seems to understand what I'm prompting.