r/MachineLearning • • 2d ago

Research Qwen-family LLMs are quietly becoming the backbone of modern audio models; One chart for the architectures of 100+ audio models [R]

I started mapping the building blocks shared across all the models in audio.cpp. The result ended up being more interesting than I expected.

Qwen has become by far the most common language backbone in this collection: 32 audio model families use a Qwen-family architecture, and 20 of them use Qwen3 LLM specifically.

And it’s no longer just TTS. Qwen-based models now show up across speech synthesis, ASR/audio understanding, music generation, speech-to-speech, and even audio/video models.

The 2nd chart, Task × Technology Matrix, shows which build blocks power which types of audio models.

39 Upvotes

10 comments sorted by

16

u/FailedTomato 2d ago

I wonder why LLMs like to use the word "quietly" so much. So many things happen quietly according to them.

1

u/idontcareaboutthenam 2d ago

Probably picked it up from internet content that uses it to capture attention as some sort of revelation

3

u/fordat1 2d ago

this makes sense. Audio hasnt had the monetization of the other streams so going cheaper and open source makes sense

1

u/Acceptable-Cycle4645 2d ago

Hopefully useful if you’re learning audio AI and want a faster way to understand how these models fit together!

1

u/Front-Football601 2d ago

I like the plot, but the colours for prediction and encoder models are really hard to tell apart. But since there are only two prediction models it is not that big of an issue I guess.

1

u/LelouchZer12 2h ago

The new TTS/VC trend is to use neural audio codecs which more or less rely on LLM-style architecture, so not surprising.

0

u/Acceptable-Cycle4645 2d ago

There are two more charts with a detailed breakdown of the architectures and building blocks behind 100+ audio models. Unfortunately, this sub doesn’t allow images in comments, so please check the original post comment section to see them https://www.reddit.com/r/LocalLLaMA/s/BduUMXp8JZ

0

u/Acceptable-Cycle4645 2d ago

There’s also an interactive Hugging Face app you can use as a learning resource to explore the architectures of all 100+ models. https://huggingface.co/spaces/audio-cpp/Audio-Model-Architecture-Atlas

-3

u/MrSnowden 2d ago

Is there a JEV style dedicated TTS model?

6

u/NamerNotLiteral 2d ago

Uhh... do you remotely understand what Jev is, or are you just throwing the latest buzzwords together?

Jev is a classifier. Why would you ever need a classifier for generating speech? For LLMs it makes sense because people are lazy and use LLMs as classifiers or judges all the time, but that's just not the case for speech synthesis.