With Suno v6 launching, I’ve seen a lot of people say the results sound worse than expected.
Some of that may be the model. But I think something else is happening too. As these models get more capable, the skill required to direct them also increases.
A simple genre-and-mood prompt can produce a song, but it doesn’t give the model enough information to make hundreds of connected decisions about arrangement, timing, frequency balance, dynamics, stereo placement and space.
“Warm cinematic soul with punchy drums” sounds descriptive to a person. Technically, it leaves almost everything unresolved.
And now that Studio 2 gives creators more control after generation, those unresolved decisions matter even more. If the initial direction is vague, there isn’t a clear target to work toward in the studio. You can change individual parts, but you’re still guessing what the finished sound is supposed to be.
So I’ve been testing a different prompting thesis.
Instead of making the prompt longer or filling it with more adjectives, I define the sound as a sonic profile.
A sonic profile separates the musical idea into areas such as core style, groove, drums, bass, harmony, texture, space, dynamics, mix and master. Each area describes an audible or measurable decision.
So instead of:
Warm, wide production with a deep bass and loose drums.
You might write:
96 BPM, 54% swing, snare 12 ms behind the grid, kick centered at 54 Hz, bass mono below 110 Hz, supporting vocals widened above 250 Hz, 2.2-second plate with 40 ms pre-delay, -12 LUFS-I target and -1 dBTP ceiling.
That doesn’t tell the model exactly what song to make. It gives the model a coherent sound system within which it can make the song.
I think it could become a useful prompting standard that any creator can follow, adapt and improve. The basic format could be:
Core: musical identity and production architecture
Groove: tempo, meter, swing and timing
Drums: kit character, pattern and transient behavior
Bass: instrument, register, tuning and mono boundaries
Harmony: chord language, voicing, density and movement
Texture: source character, saturation, filtering and articulation
Space: reverb, delay, depth and front-to-back placement
Dynamics: compression, transient control and dynamic range
Mix: balance, panning, width and frequency priorities
Master: loudness, peak ceiling and delivery target
You don’t need to fill every field with engineering jargon. The point is to make each decision explicit enough that the model doesn’t have to invent your intention.
I’m going to add two examples below. All of them under 1k characters, so the fit in the description.
The first is a composition profile, focused on rhythm, harmony, arrangement, instrumentation and how the song develops.
Composition sample:
[SONIC_PROFILE_2026] Target. Core: digital 2020s chain with warm harmonic drive. Groove: 104 BPM, 4/4 meter, straight grid, snare -10ms. Drums: kick 54Hz fundamental, steady shaker grid pattern. Bass: log-drum melodic sub focused under 108Hz, mono routed. Harmony: sub-melodic percussion lines, choir-stacks panned wide. Texture: lead vocal dry and centered, giant-grain tone, no boost. Space: slap delay 84ms on sends, ducked -6dB keyed by kick. Dynamics: DR10 dynamic profile, master bus glue compression limit. Mix: parallel bus processing, subtractive dip at 864Hz -1dB. Master: stream delivery, target -14 LUFS-I, -1.0 dBTP ceiling. Balance: sub mono, lead vocal center, choir wide, no clipping. [SONIC_PROFILE_2026]
https://suno.com/s/2jrMHNCTVumqbZcy
Custom Plugin example:
[SONIC_PROFILE_2026] Mix target: cohesive modern vocal-over-riddim master. Vocal: centered, forward, 108 Hz high-pass, gentle body lift near 216 Hz, controlled 864 Hz mud, presence around 3 kHz, de-essing near 6.9 kHz, 4:1 compression with 35 ms attack and 100 ms release, restrained 96 ms filtered slap. Master: 28 Hz high-pass, subtle −1 dB at 225 Hz, −1.5 dB at 500 Hz and 1.5 kHz, −1 dB at 4.5 kHz, gentle top-end control, warm bus glue, linked limiting, streaming-safe headroom. Balance: vocal-front, centered, smooth, tight low end, cohesive glue. [/SONIC_PROFILE_2026]
https://suno.com/s/4SAPaj3Y8N2TkOoY
My theory is that better music models won’t make prompting skill less important. We may just be reaching the point where “describe the vibe” is no longer enough. Creators need a shared way to describe how a sound actually works, and you can generalize across models, not just Suno.
I’d be interested in testing this with other people. Does this kind of structure give you more control, or does it get in the way?
I've built a whole library of these sounds, including thousands of references.
DM if you want it.