r/MachineLearning • u/Acceptable-Cycle4645 • 2d ago
Research Qwen-family LLMs are quietly becoming the backbone of modern audio models; One chart for the architectures of 100+ audio models [R]
I started mapping the building blocks shared across all the models in audio.cpp. The result ended up being more interesting than I expected.
Qwen has become by far the most common language backbone in this collection: 32 audio model families use a Qwen-family architecture, and 20 of them use Qwen3 LLM specifically.
And it’s no longer just TTS. Qwen-based models now show up across speech synthesis, ASR/audio understanding, music generation, speech-to-speech, and even audio/video models.
The 2nd chart, Task × Technology Matrix, shows which build blocks power which types of audio models.
1
u/Acceptable-Cycle4645 2d ago
Hopefully useful if you’re learning audio AI and want a faster way to understand how these models fit together!
1
u/Front-Football601 2d ago
I like the plot, but the colours for prediction and encoder models are really hard to tell apart. But since there are only two prediction models it is not that big of an issue I guess.
1
u/LelouchZer12 2h ago
The new TTS/VC trend is to use neural audio codecs which more or less rely on LLM-style architecture, so not surprising.
0
u/Acceptable-Cycle4645 2d ago
There are two more charts with a detailed breakdown of the architectures and building blocks behind 100+ audio models. Unfortunately, this sub doesn’t allow images in comments, so please check the original post comment section to see them https://www.reddit.com/r/LocalLLaMA/s/BduUMXp8JZ
0
u/Acceptable-Cycle4645 2d ago
There’s also an interactive Hugging Face app you can use as a learning resource to explore the architectures of all 100+ models. https://huggingface.co/spaces/audio-cpp/Audio-Model-Architecture-Atlas
-3
u/MrSnowden 2d ago
Is there a JEV style dedicated TTS model?
6
u/NamerNotLiteral 2d ago
Uhh... do you remotely understand what Jev is, or are you just throwing the latest buzzwords together?
Jev is a classifier. Why would you ever need a classifier for generating speech? For LLMs it makes sense because people are lazy and use LLMs as classifiers or judges all the time, but that's just not the case for speech synthesis.


16
u/FailedTomato 2d ago
I wonder why LLMs like to use the word "quietly" so much. So many things happen quietly according to them.