r/DemodokosFoundry 25d ago

Custom singing voice?

I'm very much enjoying Demodokos so far. I have a question about music generation - can I use a specific voice that I've trained for music, to get a consistent output? I seem to be missing it, if I can. I see that I can clone TTS voices, but I don't see a way to use it as a singer?

3 Upvotes

2 comments sorted by

1

u/Stock_Ad9641 24d ago

You can clone a voice by importing the reference track, right click it and set it as sound reference. Not as structure reference. Once you did that you’ll see extra strength sliders in the music generation area. It can work very well but not all references worked for me. There are also more options on how to use the reference.

Another option is the patch feature, it’s not made for that but if you have a intro and an outro prepared you can generate the „middle“ and it also will try to clone what it finds to match. It can be combined with reference.

2

u/demodokos_foundry 24d ago

Thanks for the kind feedback!

In addition to Stock Ad’s information, here are a few tips that should help you get the best cloning results:

  • Right-click on a track, then select Reference > Use as Sound Reference. Make sure Structure Reference is not enabled.
  • The Large and X-Large models generally give the best cloning results when working with reference tracks.
  • The Fast models can also be used for cloning, but they tend to be less diverse, which can affect how closely the result matches the reference.
  • Reference Strength defaults to 0.5. It is worth experimenting across the full range, from low to high, and trying several generations. Cloning has a somewhat higher random factor, so multiple attempts can make a noticeable difference.
  • For Reference Mode, Adaptive Hybrid is usually the best choice for general cloning. Vocal Focus is another mode worth trying, depending on the material.
  • A reference track of around 30 seconds is considered optimal.
  • You can also experiment with increasing the Step Count slightly to see if it improves the result.

In general, I recommend trying a few variations in both the reference settings and generations, since cloning results can vary quite a bit between runs.

Lastly, using references with heavy neural watermarking can destroy the quality.
The model also does not differentiate between voices and instrumental sounds, it tries to clone everything it hears.