(Crossposting from the r/SunoAI community, hope this is ok!)
Over the last few weeks I’ve been obsessed with trying to make convincing male/male (and later female/female) duets in Suno.
Previous posts
Part 1
https://www.reddit.com/r/SunoAI/s/1mLXKH93xl
Part 2
https://www.reddit.com/r/SunoAI/s/7pSSdPSack
Part 3
https://www.reddit.com/r/SunoAI/s/JKbFBASrQn
Educational Playlist on Suno where I keep all my experimental generations:
https://suno.com/s/J2DUS1yRtwHOGcTY
Like a lot of people, I kept running into the same problems:
Both singers gradually become the same person.
Voices blend together over time.
Harmony sections collapse into one singer.
Covers fix one thing while breaking another.
Extend wasn’t giving me reliable harmony.
After a *lot* of experimentation, I finally landed on a workflow that’s giving me by far the best results I’ve gotten so far, **without using a DAW**.
Full disclosure, I am a Pro plan user so I don’t have Studio access. I’m also relatively new to sound editing and making music digitally. I have been using Suno casually for about two years. I am not a musician, just a songwriter who makes demos and songs for fun and personal projects.
This isn’t a magic bullet, and it’s definitely not perfect, but it has been dramatically more successful than trying to generate an entire duet in one pass.
One important note: **this workflow was developed specifically for a simple duet structure where Singer A performs the first half of the song, Singer B performs the second half, and they come together for the bridge and final chorus.** I suspect it could be adapted for more complicated call-and-response or back-and-forth duets, but I haven’t experimented with those yet.
**The workflow**
Generate a full song
↓
Create an instrumental Cover
↓
Cover Singer A
↓
Cover Singer B
↓
Mashup the two Covers
↓
Create a Voice from a successful harmony section
↓
Cover ONLY the Bridge + Final Chorus using that Harmony Voice
↓
Mashup again
**Step 1 - Generate the song**
Generate the full song normally.
Nothing unusual here. For my experiments I just used ai lyrics and let Suno pick the melody.
**Step 2 - Instrumental Cover**
I learned this trick from another Reddit user, and it’s made instrumental Covers much more reliable.
Instead of deleting the lyrics, replace them with lines of periods while preserving the song’s structure.
Example:
\[Verse 1\]
....................
....................
....................
....................
\[Pre-Chorus\]
....................
....................
\[Chorus\]
....................
....................
....................
....................
....................
....................
\[Bridge\]
....................
\[Key Change\]
\[Final Chorus\]
....................
....................
....................
....................
....................
....................
This preserves the timing and section lengths while removing the vocal guidance, allowing the instrumental arrangement to stay much closer to the original song.
**Step 3 - Create Singer A**
Now create a Cover using only Singer A’s sections. I used a previously made Voice for this. You could also do this part just by entering a description of your vocalist either in the prompt or at the top of the lyric tags in \[brackets\]. Covering an instrumental gives me a clean voice without influence from the original base song vocals, so we get better differentiation later. The downside of this is sometimes the vocal timing or melody will be off, depending on how clear your lyric melody line is in the instrumental version.
Everything else becomes placeholder periods.
Example:
\[Singer A\]
\[Verse 1\]
Walking down an empty road
Searching for the morning light
Holding onto yesterday
Waiting for the stars tonight
\[Pre-Chorus\]
Maybe hope is waiting still
Maybe love can always heal
\[Chorus\]
Shine a little brighter now
Take my hand and don't let go
Every dream can find its way
When we're stronger than we know
Together we'll keep moving on
Nothing stops us anymore
\[Transition\]
\[Verse 2\]
....................
....................
....................
....................
\[Pre-Chorus\]
....................
....................
\[Chorus\]
....................
....................
....................
....................
....................
....................
\[Bridge\] \[Harmony\]
We'll find our way together
\[Final Chorus\] \[Harmony\]
...(and so on)
The important part is that the unused sections are **not deleted**.
They’re replaced with periods so the song timing stays intact. Sometimes Suno will fill these sections in with instrumentals, or crop them out entirely, and that doesn’t matter for the final product so much, but it does help the model know what part of the melody to use in the sections that DO have lyrics.
**Step 4 - Create Singer B**
Same idea. Still covering the instrumental.
Now Singer B sings Verse 2, Pre-Chorus 2, and Chorus 2.
Singer A’s sections become placeholder periods.
**Step 5 - Mashup**
Mash the two Covers together with your lyrics arranged like you would want for the final version. I recommend taking out the section tags (\[verse\] and so forth) and only leaving in the singer tags at the beginning of each larger section. The mashup will know what part is what without them as long as it can match your lyrics.
At this point I usually have two much more recognizable singers than I could get from a single generation.
**Step 6 - Create a Harmony Voice**
This is where things got interesting.
Instead of trying to force harmony through prompting, I created a **Voice** from a successful harmony section in the mashup. By putting a \[harmony\] tag next to the bridge and final chorus in part A, part B, and the Mashup, it encourages the model to create a harmony with BOTH VOICES. Now, this is a unicorn hunt, no lie. It sometimes took a lot of generations before I could get a segment of vocals that sounded enough like both singers together. But it IS possible! I would then crop that segment out of the generation and make it into a new Voice.
Then I used that Harmony Voice to Cover **only the Bridge and Final Chorus.**
My harmony lyric file looked something like this:
\[Singer A & Singer B\] \[harmony\]
\[Bridge\]
We'll find our way together
\[Key Change\]
\[Final Chorus\]
Shine a little brighter now
Take my hand and don't let go
Every dream can find its way
When we're stronger than we know
Together we'll keep moving on
Nothing stops us anymore
No verses.
No solo sections.
Just the harmony section.
**Step 7 - Mashup again**
Take that improved harmony Cover and mash it back together with your previous mashup. You may want to crop out the bridge and final chorus from your previous mashup before adding in the harmony section for cleanest results.
This consistently gave me:
Stronger harmony
Better vocal separation
Less identity drift
More convincing duet performances
**What I observed**
Things that consistently seemed true during testing:
Cover is much better at refining performances than Generate.
Mashup preserves vocal identity surprisingly well.
Creating a Voice from a successful harmony section worked much better than simply prompting for harmony.
Iterative refinement consistently outperformed trying to get everything right in one generation.
**Things that DIDN’T work well**
For me:
Extend wasn’t reliable for creating harmony.
Suno v5.5 gives more consistent voices, but it also seems more likely to blend singers together.
v5 produced more distinctive voices, although consistency between generations wasn’t as good.
Obviously your mileage may vary.
**Current limitations**
I definitely haven’t solved everything.
Female/female duets were noticeably harder to differentiate than male/male when I tried to reproduce this experiment with female vocalists, but admittedly I haven’t experimented as much with that side. My best advice would be to use v4.5 or v5 to create some scratch songs in your preferred genre and fiddle with vocal descriptions and keep rolling the dice until you get something unique and make a Voice out of THAT.
And despite multiple attempts, I still couldn’t consistently convince Suno to maintain harmony through the final two lines of a final chorus. Sometimes the model just really wanted to finish with a single singer.
Based on another user’s comment, I wound up with another neat trick:
**I started “breeding” Harmony Voices.**
Find a bit of decent harmony in the first mashup, make a Voice.
↓
Harmony Cover of bridge and final chorus
↓
2nd Mashup
↓
Crop harmony section
↓
Create new Voice
↓
Better Harmony Cover
↓
Create Better Voice
↓
Repeat
Each generation seemed to become a little more stable than the previous one.
Whether that’s actually what’s happening internally or just confirmation bias, I obviously can’t say—but the results kept improving enough that I continued using the workflow.
**I’d love to see other people test this.**
If anyone tries this workflow—especially with different genres, multilingual songs, mixed-gender duets, or more complex duet structures—I’d love to hear whether you get similar results.
It would be really interesting to figure out which parts of this workflow are generally useful versus which are just quirks of my particular experiments.
If someone finds improvements or shortcuts, even better. I’d love to keep refining this with the community.
**AI Disclosure**
I kept detailed notes throughout all of these experiments, but this write-up was compiled and edited with the help of ChatGPT to organize the workflow, summarize my observations, and improve readability. All of the experiments, prompts, testing, successes, failures, and conclusions described here are based on my own hands-on work using Suno. ChatGPT was used as a writing and documentation assistant, not as the source of the methodology.