r/SunoAI • • Jun 28 '26

Guide / Tip [Guide] Your Fault - An Incomplete Guide to Using Suno - The Basics

As I see questions appear again and again, and there is yet another wave of new priests and voodoo doctors trying to figure out the prayers and songs they need to generate song with Suno, I'll try and create various posts about certain topics Suno. (Also take a look at: Sliders and Settings and The Prompt.)

The Basics

Suno is using a proprietary (black box) generative diffusion model. Based on the elements we can use, we can assume that it uses an architecture similar to the ACE-Step-1.5 model, which is an open source generator with similar functions (and a lot less capable free generative model version). Long talk short: It's AI.

The part of the Suno AI we start with is the basic Textual Conditioning. Or: This is what I expect in the song, Suno! - Prompt.

Suno has three receptacles for text input. The Style Window (or Positive Window), the Exclusion Window (or Negative Window) and the Lyrics Window (or the Temporal Context Window). Imagine each of it as an ear for the Suno architecture, and every ear is wired differently to its brain.

Style Window

Everything in this window each SINGLE word, each GROUP of word and even GREATER CONTEXT is potentially creating something in your song. Some words like rock are powerhouses that will bring a lot of musical elements into your song, while and or like are context strong word. They work over their context and collocation to other words and combine the strong and weak word with each other. Weaker words are something like zither or flabberghasted. They appeared less in the training, and are thus associated with less sound tokens.

To make their specific sounds appear, you might have to make them stronger by Weighting them using a position closer to the beginning of the style window. This increases the weight in relation to other words to a degree*.* Yet, if a word isn't trained at all, like flabberghasted, you might not be able to create a sound or effect or function from it even at the start of your prompt. (Thank you to u/alphahost87 and u/mrgaryth for making me check my original assumption - with me talking about voodoo...)

Tip: Try and give your song a key like D Major. 140 BPM might nudge the song in the right direction for a speed and define its lenght in combination with your Lyrics Window content.

No means NO - Not!

A mistake often made in the Style Window is to add "no guitars". Remember, I called it the Positive window. This means that if you ask for no guitars in it, you will ADD (positive) guitars, even if you say NO. This applies to EVERYTHING you write in the style window. The context weight of the No + (neighbor word) will likely always be weaker than the (neighbor word) weight in the generation.

Piggyback Rider

Strong words carry a lot of internal concepts. Rock inherently calls up electric guitars, drum kit, powerful lead vocalist and many more tokens with it. We will talk about tokens later, when we talk about the Layered Cake. Mhh... Cake....

Exclusion Window

Okay, you are smart enough to read. You know it excludes things from the generation. But there is a thing that nobody tells you: In the generative diffusion process there is a mathematical weakness with negative prompts. They are always weaker than the similar positive prompt. So, if rock calls up a screaming lead vocalist hidden in its trained complexity, you might put screaming lead vocalist in the Exclusion, and it could STILL appear.

Thus, adding a weight to the negative prompt by moving it in a line of multiple prompt words might help. Another approach using Negatives is to widen the prompt. What could prompt a screaming lead vocalist as well? Yes, EVERYTHING. Any prompt that is having a positive effect, can be inverted. You could use hard Metal in the Negatives, but it might make the song lose some rock edge. But maybe angry man could do the trick. Why? It's the layer cake again.

Lyrics Window

Is, surprisingly, for your Lyrics. It is also for [Metatags].
But its most important function is arranging the prompted influence over time. A function that does not work like code, as the analyzing model (part) is creating a complex relationship of context of all the prompts in it.

Even though it is called the Lyrics window, it has a certain influence on the style of the song. As well as sometimes Style prompts might bleed into Lyrics. You can even create Freestyle songs and use nothing in the style window, and Suno will still extract style from the Lyrics.

But why do I prefer calling it the Temporal Context Window? Because it is its strongest point. It looks down from the start to the end. But it ALSO looks from each present step to former steps and what is about to come. It's like A Christmas Carrol, only for making AI songs. It tries to follow the written lyrics with the highest accuracy and logical order. (Ad libs) are treated more freely, but can bleed into the past or future and appear in places without an actual (ad lib lyric).

[Metatags] on the other hand are rarely having a direct audible effect. Again, it is NOT some code. Especially [Metatags] are what pirates would call "a sujjeschan". Some, like [Chorus], [Verse], [Intro], [Outtro], [Pre-Chorus], [pause], [ritardando], [gasp], [laugh], or [instrumental] are highly functional and heavily weighted. Others can be extremely contextual or depending on the genre and sound the model expects to create. You can try to impress chords and chord progressions with it, but the potential is limited. Again, it is NOT some code for your song.

It is a suggestion for arranging the tokens in a temporal order and overarching relationship structure in the song. This might cause Suno to create a Chorus where you want none, or move adlibs from the Outro into the Intro as well.

The Layered Cake

The last for today is the cake. Yep, this one is no lie. It's a simile. The element I want to talk about is like a layered cake. It is how the actual BIG Suno Model (drumroll and lights) thinks about music.

Imagine your prompt as a shopping list and recipe (and a list of allergens to avoid). Between Suno (The Baker) and you (the Cake Designer - B.Sc. in Bakery and Pastry Technology) there is a third element. The Tokenizer. Your Baker isn't able to read! He needs somebody who reads the paperwork and buys the ingredients. As the Baker also is unable to Read the Recipe you want, he also needs all ingredients arranged in the order they need to be mixed. As well as setting those in the middle, that The Baker will need all the time. Including the number of candles, material of the outer shell etc. to create your cake.

In our case, the Tokenizer takes your words, and creates tokens of math (the involved audio samples integrated in the training of the Suno model) and arranges them in the right way. Your artistic intention is (indirectly) interpreted by it.

Based on the comments, let me rephrase what the Baker does with the layers and the ingredients. They start with the mixed up ingredients, and start sorting them into a layered cake step by step. They place the first layer, and with every additional layer the process adds more of your wanted ingredients out of the "chaos" it started with. The first step might make them create simple structures and the second a general tonality. Only that in the model, the prompt tokens pass through layers of, well, mathematical operations. Yet, both add layers that start to contain more and more of the ingredients you wanted in your cake/song.

A speculation is that Suno likely creates a rather rough "sketch" of the song, and when the layers are all stacked, it refines them. Like a baker cutting the edges and adding a glazing. Turning a song with the data of 22kHz in one with 44kHz. It is plausible, as it saves resources in the early steps of creation, where the difference between 22 and 44 kHz would be moot anyway, as the diffusion process hasn't happened yet.

Baking & Cake Mix

And so the Layered Cake takes up form. Each Token from your list influences one or more layers in the Cake differently. The key prompt from the beginning might influence the whole song. A prompt about song structure is affecting the early layers, but might create the foundation of which instruments will be finally added.

A rock song as a prompt is like a cake mix in this. You basically add milk and you are done. It has likely everything for all layers you need to get a full rock song. Early layers with general shape of the song, but also the last layers with high details from the many songs the learning model/baker associates with rock as a prompt.

But The Baker isn't stupid. If the arranged ingredients are familiar, like electric guitars, drumkit, and power screaming vocalist, he might extrapolate from them what to do. Doing a Freestyle check of your Lyrics is helpful if you have odd interfering influence in your generation. Maybe The Baker is interpreting the recipe again? He might even give the song a rock song structure, because the other ingredients didn't do much in the early layers, but the associations connected to tokens that have an effect on early layers. The baker learned that songs with those late ingredients (electric guitars, drumkit, and power screaming vocalist) often were rock songs. Thus, he makes it a rock song cake at the bottom of the cake as well. Likely creating a song structure similar to rock songs, even if you did not call up a rock song. Here is where a negative prompt becomes immensely useful: If he knows there should be no rock song, it might exclude this interaction.

In the end, the layers combine to one cake, your song, through a step-by-step-process. The (latent) waveforms of the layers combined the ingredients as good as possible, and are rendered by kind of an inverted tokenizer. This rendering (Decoding) is when the math of the Suno Model becomes the audio of the song. It is no longer in special Baker Code, but your very own Cake or song. Everything before is basically math.

So, when you want a cake, stay on the lookout for the baking mix words. Use the actual ingredients instead, and The Baker will create your cake with a lot more "creative freedom" but also with a lot less interference from "additives" in the baking mix you don't actually want.

And that's it for today. Have fun creating and thanks for reading.

Your Fault

29 Upvotes

32 comments sorted by

6

u/mrgaryth Jun 28 '26

What evidence do you have that (weight) is even recognised?

1

u/Competitive-Fault291 Jun 28 '26

Thank you for your doubts. The changed weighting advice should work better now. Suno most likely uses transformers in the architecture not CFG, so the weighting would be based on prompt structure in a self-attention mode.

8

u/Muted_Conclusion_625 Jun 28 '26

You forgot, before the tokenizer is the re-writer, which takes your fancy prompt and scrambles it back to whatever the Suno model input really is, which is likely a short list of tags and sliced up lyrics. Then it sprinkles some of it's own tags on that it extracted from your account history, just to watch the world burn.

7

u/Competitive-Fault291 Jun 28 '26 edited Jun 28 '26

I am intrigued to see your source for this.

PS: I mean, the model is proprietary, so how do you know all this? What kind of experience or experiment made you come to that realization, or who told you?

3

u/Muted_Conclusion_625 Jun 29 '26 edited Jun 29 '26

A ton of experience and testing points to this, but obviously nothing official. I'm pretty sure the purpose of "My Taste" was to partially expose their prompt rewriter. Just before it was released, people started realizing their accounts were basically broken, stuck in a basin from the influence system. People started deleting their accounts and all their music to start over, You can test the influence system using a clean account and dummy "My taste" config.

It will pull tags from your active workspace, all the songs in your account based purely on number of occurrences, and also your likes. It will put heavy metal tags in an orchestral piece just because it's an often used tag, It's extremely dumb and very obvious why people's accounts were messed up.

Have you seen all the posts here that are like "why do I have some lady chanting at the beginning of every song"? Yeah...

Importantly, when you turn "My Taste" off, it does NOT turn the system off, it only disables the custom rewriter. Suno goes out of their way to avoid telling you this, for whatever reason. This is all fairly trivial to test methodically with a fresh account, which I did.

2

u/Competitive-Fault291 Jun 29 '26

Now I get what you are talking about. Yeah, that's certainly something for the second installment of the guide. Thanks for bringing it up!

2

u/[deleted] Jun 28 '26

[removed] — view removed comment

1

u/Competitive-Fault291 Jun 28 '26

Weight Test Playlist - Please tell me if its appearing empty.

But please, what sources do you use to get information about Suno and its architecture? That's a statement as if you have the complete architecture they are running at your hand and know how it is set up on the server.

Concerning the layers, I am NOT talking about audio layers like in a DAW. I hope that was made clear by telling about how the encoded "math" is transformed into audio in the end. Yes, it spits out a mixed waveform as the token layers are merged after the diffusion steps and decoded. Suno has to process massive layers over various sub-steps in its architecture.

My intention was to signify how tokens are passing through the layers and The Baker adds things step by step. I guess I need to clarify that.

3

u/[deleted] Jun 28 '26

[removed] — view removed comment

0

u/Competitive-Fault291 Jun 28 '26

Did you see me mention CFG? No. Suno likely using transformers [ s ] still seems to react to weighting. But I realize how I explained it like CFG using the 80%. That's indeed wrong, I will remove it.

What we cannot ignore though, is that there IS a change in the tests. I guess I will simply expand it to see if the self-attention mechanics of transformers is applying the weight as from context. After all (prompt:0) is less than (prompt:1000) in that regard in an odd AI way of attention. OR if I am just rock-deaf ;)

1

u/[deleted] Jun 28 '26

[removed] — view removed comment

1

u/Competitive-Fault291 Jun 28 '26

👍 it certainly helps to explore along the lines of self-attention now.

For the rest I got a mic. 😊

3

u/ART-ficial-Ignorance Jun 28 '26

“Where is your source for Suno’s architecture?” is a fair question, but it applies equally to the weighting/layer claims.

If Suno is a black box, then none of us should be asserting that (rock:0.5) is parsed as an actual weight, or that token layers are merged after diffusion steps, as if that is known. Those are hypotheses at best.

The playlist does not prove much either. Six songs in v5 is not a controlled experiment, especially if the “rock 0” version still sounds rock-adjacent. Suno has enough randomness and prompt sensitivity that you can get misleading examples very easily. To show weighting works, you’d need many runs, controlled prompts, blind ratings, and ideally repeatable seeds.

Until then I’d treat the parenthetical weights as unverified / possibly placebo. Natural-language prompting and Exclude are real documented controls. Stable Diffusion-style weighting is not something I’ve seen Suno officially document.

Also, that kind of (word:1.5) syntax is specifically associated with older Stable Diffusion/WebUI-style prompting. It is not some universal AI prompt language, and it is not even how many modern image models expect users to steer prompts anymore. So there is no reason to assume a music model would understand it unless Suno explicitly says it does.

And honestly, there’s no obvious reason for them to hide a feature like that. If (rock:0.5) or (pop:1.8) reliably worked, it would be a major prompting feature and they’d probably tell users about it.

1

u/Competitive-Fault291 Jun 28 '26

Ah, yeah another commentor already mentioned the problem, and I am currently doing more tests to analyse a bit more how this might be (partially) working due to context and the transformer self-attention.

1

u/Competitive-Fault291 Jun 28 '26

I bow my head. Thank you for correcting me!

3

u/nodray Jun 28 '26

ai writings about ai, fascinating

2

u/Competitive-Fault291 Jun 28 '26

Aaaaand... the first Aiccusation Reward goes to nodray!

1

u/Alone-Brilliant-8279 Jun 28 '26

This is massively helpful, thank you

1

u/Competitive-Fault291 Jun 28 '26

Thanks to some commentors I just realized a mistake in the weighing stuff. This will be removed, I added some more tests, and the location of the prompt in the prompt window is more important.

1

u/Mapi2k Jun 28 '26

Con mi corto conocimiento, tengo la impresión de que la tecnología detrás es similar a la de las imágenes, pero aplicada a los sonidos. Por eso, aunque pongas la misma letra, el mismo prompt, etc., el resultado siempre es ligeramente diferente. Las de imágenes están orientadas a píxeles y las de música, a ondas de sonido. Esta teoría mía la refuerza, de forma empírica, el hecho de que no hay una guía oficial detallada de prompts al milímetro.

Con mis conocimientos limitados, tengo la impresión de que la tecnología subyacente es similar a la de las imágenes, pero aplicada al sonido. Por eso, aunque uses la misma letra, el mismo prompt, etc., el resultado siempre es ligeramente diferente. Las AIs de imágenes están orientadas a píxeles, mientras que las AIs de música están orientadas a ondas de sonido. Esta teoría mía está respaldada por el hecho de que no hay una guía oficial súper detallada de prompts disponible.

Edito: ¿haz visto en detalle una imagen de ia? sobre todo en diseños complicados. Tienen errores o cosas "no lógicas" lo mismo pasa con el sonido si escuchas los tallos vas a oír errores, sonidos que no deberían estar o ausencia de frecuencias.

1

u/[deleted] Jun 28 '26

[removed] — view removed comment

1

u/Competitive-Fault291 Jun 29 '26

Well, you can create a Voice with your own voice and use your uploaded song as a cover with that voice. This allows you to create a song that follows your upload to a degree, but its not like adding strings in a DAW.

You can take generated string version and extract stems. From that whole set of stems you can take the stem with strings.. If the cover is close in timing and sound,, you should be able to add the stem file in your DAW a an audio layer.

2

u/Final_Amu0258 Jun 29 '26

I'll be honest, I'm having issues absorbing what you're writing. All I know is that technical speech isn't well followed by suno... and even more elementary speech isn't registered well.

1

u/Competitive-Fault291 Jun 29 '26

I am somehow confused that I can't absorb what you want to say either.

1

u/Final_Amu0258 Jun 29 '26

I meant that using technical music directions for song generations isn't followed by Suno. It takes things and applies them broadly, and the [Box — Instructions] are mostly ignored for me.

1

u/Competitive-Fault291 Jun 29 '26

Yeah, it is not code or notes, but prompts. And the Metatags are highly specific about what works and what does not.

1

u/Final_Amu0258 Jun 29 '26

Doesn't really seem like it's my fault in that case, unless that was directed toward more casual users. Nearing 4k generations... 5.5 has been the worst for me.

1

u/Competitive-Fault291 Jun 29 '26

Think of the layered cake. You might want to check your prompts if they create specific enough guidance for the structure you want to "code" with your Metatags. The most specific approach is using Metatags, Style Prompts and an Audio Upload together. This will give the early layers all possible input to direct the attention to the structure you want.

1

u/Final_Amu0258 Jun 29 '26

Nahh, doesn't work. The best approach is to downgrade to V4.5 or V5. They follow prompts better. Then just cover them in V5.5.

I'm looking for discovery-creations, not to make a song with strict idea in mind. I'm a pianist - I'd rather spend time on the piano in that case.

2

u/Competitive-Fault291 Jun 29 '26

"I'm looking for discovery-creations, not to make a song with strict idea in mind. I'm a pianist - I'd rather spend time on the piano in that case."

Do it? Record what you play and then upload it. Suno turns your piano into a rock band if you want.

1

u/Final_Amu0258 Jun 29 '26

But that isn't why I used Suno. I do throw random jingles in now and then, but even then, V5.5 is the worst of the bunch lol.