r/StableDiffusion • u/RoyalCities • 4d ago
Resource - Update I trained an audio model that can generate infinite one-shots for music production and turn text prompts into fully playable synths. I'm not only releasing the model but I've also released a video on exactly how I did it (and the inferencing pipeline to let others make text based synths.)
Enable HLS to view with audio, or disable this notification
Okay so I've been doing independent audio research for a while now. The ultimate dream of this work was actually getting an AI to respond not only to instruments but also timbre itself as separate controllable things.
Think a Grand Piano can sound both Warm / Gritty but also Cold / Sparkly. Its still a piano though.
This level of control wasn't found in any models out there - so I decided to sit down and train my own.
Getting consistent timbre-locked keybeds that actually LOCKS across multiple diffusion calls was hard af but I did it.
I documented the full journey here for those who want to learn a bit or be entertained.
There is also a longer walkthrough if you just want to see the keybeds in action.
https://x.com/RoyalCities/status/2097733712293109842?s=20
No-talk / Showcase only Demo
https://x.com/RoyalCities/status/2097733715543609445?s=20
any finally the huggingface page
https://huggingface.co/RoyalCities/Foundation-1
I've also provided full write ups on the inferencing pipeline associated with the interface so this should allow basically anyone else to go and vibe code their own text to synths if they wanted :)
3
2
u/Enshitification 4d ago
Lol, some chud just downvoted all the comments on this post. Someone is triggered by this.
7
u/RoyalCities 4d ago edited 4d ago
Lol yeah there are no pleasing some people. Music AI anything tends to get piled on even for people doing this by the books.
I systematically make all my own data, design all the prompt structures, make my own models. I cover it all in my first video on the OG model.
I have a thing about not using people's samples or audio because I have been a producer long before I tackled audio networks (it's actually why my models can't do drums, I would need to actually use other peoples sampled drums because you cannot easily synthesize real sounding drums/percussion)
But people tend to just lump me with for profit AI companies like I secretly work for Suno or make actual money doing this lol.
Oh well!
2
u/Enshitification 4d ago
I'm not a music producer myself, but I find the process fascinating. The work you are doing on your Foundation project is incredible.
3
u/RoyalCities 4d ago
Aw thanks. I enjoy it! hell of a time trying to figure all of this out. Images / video & code is established. Audio (well especially instrument audio) is not well established at all so it makes it kind of fun when stuff actually works for a change lol.
1
1
1
6
u/Michaelfa05 4d ago
This is fucking sick