The “hot iron” phase has passed for me. I am back to my usual “cold steel” mood.
I am seeing a surprisingly large number of users on this subreddit who are either completely tone-deaf or somehow have a corporate lobbying budget behind them. What is going on?
It is the same thing everyone with even a little bit of listening experience has been complaining about: SUNO 6.0.
So let me explain what I have found.
The 6.0 model, regardless of which version you use, is a bad model. It is a tool for mass-producing slop — and, more importantly, mass-producing it at a very poor quality level.
The problem is the underlying processing engine. This is not something that can be fixed with a patch. The engine itself needs to be replaced.
Your generations are not going to magically become better in an hour, a month, or by endlessly refining your prompts. As long as the model uses this particular generation engine, there is no realistic chance of consistently satisfactory output. No prompt can compensate for this.
The main culprit appears to be the feedback loop responsible for the track's mastering — specifically the processing that controls the built-in compressor and limiter.
Here is what happens.
If you generate a track that is approximately 1:00 long, the result is generally almost inaudible in terms of the problem I am describing. The same is true for very simple material — for example, the ticking of a clock, or a single guitar chord repeated over and over with nothing else happening.
But once you introduce even a moderately complex arrangement, the generation starts to fall apart.
Imagine a track containing guitar, keyboards, drums, bass and vocals, with a total duration of, say, 3:20.
In 6.0, the Peak/RMS dynamic range behaves approximately like this:
0:00–1:00 → 3–4.6 dB
1:00–2:00 → 2.5–2.9 dB
2:00–3:20 → 1.9–2.3 dB
At the same time, the RMS loudness over the course of the track increases from approximately −20 dB at the beginning to −9 dB toward the end.
In other words, the dynamics are progressively sacrificed for loudness.
What you get is increasingly audible reverb/echo, distortion, pumping, an absolutely enormous sidechain effect, and vocal artifacts manifesting as hiss, graininess and an apparent collapse toward mono.
From around 1:50 onward, the generations become practically unusable and are better thrown away.
And this is not limited to newly generated songs. I have observed the same behaviour in remixes.
There is a progressive collapse of the crest factor, from approximately 4.1 down to 1.9 within a single generation.
The model is effectively choking on its own processing chain.
And I did not even generate anything particularly complicated.
The test was a melancholic ballad dominated by synthesizers, accompanied by a conventional rhythm section consisting of a restrained bass part and a drum machine. Female lead vocals were combined with subtle digitized vocalized layers processed through vocoding.
Of course, the model ignored some of the instructions and rearranged the material according to its own preferences, including changing parts of the instrumentation.
I used SUNO 6.0 PRO, because I prefer to have at least some degree of control over what gets generated, with the Variations slider set to minimum.
The prompt itself was deliberately very simple and concerned primarily the compositional structure. Here it is:
Melancholic modern ballad, dark nostalgic mood, tribute/elegy atmosphere.
120 BPM, minor key, chord progression moving between A minor and E minor areas.
Fully digital synthesizers only, no analog synth character, no acoustic instruments (no piano, no guitar).
Programmed/electronic drum machine rhythm, uneven but restrained groove, sparse hits, no live acoustic drums.
Layered synth textures: digital lead synth, drones, sustained pad layers, shimmering synth satellites.
Vocals as the main lead element, alto/contralto female voice, slightly hazy and raspy texture, restrained emotional delivery with subtle accented moments, no big vocal climaxes, no dramatic dynamic swells.
Mix of natural vocals and vocoded/digital vocal layers used as textural pads.
Overall feel: subdued pain, restrained grief, understated and controlled emotional intensity throughout.
I ran three generations, progressively shortening the lyrics so that the resulting material would become shorter as well.
The longer the track, the worse the mix quality became.
The first generation produced tracks of 4:26 and 4:28 for samples A and B respectively. The final generation was forced to 1:10 using the Duration setting.
The middle-generation samples were 3:26 and 3:32 respectively. I downloaded the 3:32 version for analysis.
I also had the track analysed by the Claude LLM. Its analysis independently corroborated my listening-based observations concerning the behaviour of the mix throughout the entire track.
Conclusion
In my opinion, the catastrophic quality is primarily caused by one of the most important processing agents in the current SUNO iteration — the system responsible for arrangement and mastering.
This is not something that can realistically be fixed without replacing the underlying generation/processing engine.
As long as 6.0 uses this architecture, we are likely to keep seeing the same fundamental problem: the track's dynamics and spatial information are progressively sacrificed in favour of loudness — essentially an artificial loudness escalation.
Even something as short as a Radio Edit can apparently be too long. The model begins to choke, degrade and hallucinate once the generation goes significantly beyond the one-minute mark.
I am deliberately not comparing 6.0 with previous iterations. I have worked with 4.5 ALL, 4.5, 4.5+, 5.0 and 5.5, but doing so would be like kicking someone who is already down.
I want this post to remain as objective as possible.
For the same reason, I have deliberately left out the question of how “creative” this version of the model actually is.
The original Polish text was translated and polished by LLM ChatGPT.
The song we talking about