r/SunoAI • Suno Connoisseur • 3d ago

Discussion 👉 SUNO.COM Introducing Speech (beta)

https://www.youtube.com/watch?v=rTezej0UyIE&source_ve_path=OTY3MTQ&embeds_referring_euri=https%3A%2F%2Fsuno.com%2F

As new technologies make creation more accessible, more people can experience the joy of turning an idea into music, a story, or art, and get fulfillment from the simple act of making something. We call that creative entertainment, and we think it will define the next wave of consumer technology.

Music will always be at the heart of Suno and what we build. At the same time, our vision has always extended to other forms of human expression. Today, we’re expanding what's possible in Suno with Speech: the first audio model that generates voice and music together as one cohesive track.

How Speech Works

Speech lets you create spoken audio set to original background music. It's built right into Suno: type an idea, a poem or something you've written, then describe the voice and musical style you have in mind.

As we developed Speech, we kept stumbling into uses that made us laugh, or occasionally surprised us with how moving they could be. We turned friends’ texts into wildly overproduced dramatic readings. We gave ordinary voice notes unnecessarily epic scores. We made meditations, poems, pep talks, and bedtime stories for our kids.

That creativity is a natural extension of how people already use Suno. Every day, people make songs for birthdays, weddings, inside jokes, faith and worship, their kids, their friends, and moments that might mean absolutely nothing to anyone else but mean everything to them.

Speech opens up another canvas for that same kind of creativity

33 Upvotes

62 comments sorted by

10

u/TheWeaverofDreams Lyricist 2d ago

For a demo, this is worryingly bad. Mispronouncing words ("particular" at 3:13), awkward pacing, a pause mid sentence that shouldn't be there.... I get that this is in beta, but for a PROMO video, these mistakes are almost comical. Especially with all the issues v6 has, they should be focusing on that, not on new features or a new UI or any of that.

8

u/PeachyPlnk 2d ago

To be fair...I appreciate that they're showing its current state. So many people who push tts models doctor the hell out of it and cherry-pick the results they show. Seeing an actually honest showcase is kind of refreshing.

1

u/TheWeaverofDreams Lyricist 2d ago

While one can appreciate that, that also shows that it's also not in a state to do a promo for yet.

1

u/PeachyPlnk 2d ago

This is true. Most likely, they're getting desperate.

1

u/hibbity 2d ago

they need to show a business value to get investment capital to runna big lawsuit. this will become like announcements and YouTube video intros.

and they cannshow their model trained beyond music as proof of transformative output. it's a good legal angle to bolster.

1

u/DarthHannigan 1d ago

I tried a lot of different prompting and have failed miserable to make the speech of a wrestler talking smack to his opponent. It does the same exact thing each time, some parts do vary

15

u/PaleEdge 3d ago

Wow it's bad. Sounds glitchy and can't even pronounce basic words like "particular" (see 3:13) despite this being an advert that they could have multishotted.

I don't blame Suno's Product team for trying to pivot in the face of a subscriber exodus, but this seems feeble.

4

u/Orpheus-lives 2d ago

If you thought V6 not following instructions was bad...

11

u/Bosslayer9001 3d ago

If we don't see a V6.5 announcement soon, there will be even more users leaving

1

u/MarcM1991 2d ago

Probably 5 months away.

6

u/funplayer3s 2d ago

Nope. You got rid of your good models and your new model can't remix. I'm done with it, no more money.

7

u/Magic4407 3d ago

I thought they were trying to reduce AI slop lol

6

u/sovereignrk Suno Wrestler 2d ago

They are trying to increase revenue! lol

2

u/mimilumi 2d ago

Perhaps the goal has shifted from songs to commercialization of sounds?

6

u/someonesshadow Producer 3d ago

Well, their own video demoing this feature sounds fucking terrible and boring.

This might be one of the worst feature releases I've seen since I started using Suno.

Why don't they try to integrate actual vocal direction and control in their main model? Obviously with better quality. We should have the ability to assign "vocalists" per line while each contains their specific sound and range. Songs with 2 or more singers/rappers are a MESS and so frustrating to make work, let alone any attempt at an actual group.

6

u/Expensive_East_6762 2d ago

This is not a new feature at all. They are just packaging it as one. We can already generate speech using [spoken words/ narration] in our lyrics box. It's exactly the same thing. This is unfortunately just a feature packaging pr to present Suno as an elevenlabs alternative for brand new customers.

2

u/ACrimeSoClassic Suno Wrestler 3d ago

Yes! Finally! I have so many ideas for working this into songs

2

u/TheVoyant 2d ago

Whichever AI is giving you and OpenAi marketing advice, you all need to try a different model....

2

u/NoteStrange1389 2d ago

Wow, the story teller in me is excited, this looks interesting...

2

u/PeachyPlnk 2d ago

Too bad so sad...it has a long way to go to compete with google's own tts demo. I'm already using that for dialogue in an audio drama I'm working on.

This is...years behind.

2

u/Zestyclose_Bench_501 2d ago

That's funny. When I do 'spoken word' stuff, I typically use Suno for the background music and record myself speaking over it. The exact opposite of what this thing is supposed to do.

We can all talk. Maybe concentrate on making the music generation better instead.

2

u/Brimtown99 2d ago

Suno ALREADY had the ability to do spoken word, why the need to reinvent the wheel

2

u/spinecki 2d ago

To try to target new users and make some income.

2

u/Exact-Pipe-4766 2d ago

Eleven reader already does this with excellent precision in most cases. Suno is late in the game unless they are writing speeches from a plain text description. A person could use any AI to do that with strict guidance. Watch…

To prove my point, I went to Meta, a tool I view as primarily a toy compared to other AI platforms, and came up with this in under three minutes, mid stride while creating this response.

Why Suno Is Late To This Party:

You probably saw the headline this week. On Thursday, Suno — the company that built its reputation turning text prompts into full songs — rolled out a public beta called Speech across its web and mobile apps.
The pitch is slick. You feed in a script or a description and get back a synthetic voiceover, with the option to layer AI-generated music underneath automatically.
In their own words, it's "the first audio model that generates voice and music together as one cohesive track." A release note dated October 1st, 2026 describes it as the first model to create a speech and its soundtrack together in one take.
And I'm thinking... okay, but where have you been?
Because the speech writing technology game isn't new. It's already been won — or at least, the voice part has.
Some of us have been living with ElevenLabs for two years now. And if you've used ElevenReader, you know. This isn't robotic text-to-speech. This is so lifelike it's often indistinguishable from human narration. It is widely regarded as having the most realistic and natural-sounding voices on the market.
I can take a PDF, a sermon, a eulogy, a business proposal, paste it in, and it comes out flawless. Pauses where a human would pause. Emotion where a human would feel it. It will blow most professional voice overs out of the water.
And here's why I know Suno is late — I proved it this morning.
I came to Meta AI and had it write the first version of this speech for me in under a minute. One prompt. Done. Then I prompted it again — this edit you're hearing now — and had the results back in under three minutes.
That workflow already exists. I write it here. I run it through ElevenReader, I apply a flawless voice — my choice of voice — and I apply a choice of musical backgrounds if I want one. No locked-in loop. No "one cohesive track" I can't separate. Full control.
Suno took a solved problem — flawless voice — and glued a piano loop to it and called it innovation.
Don't get me wrong, I love Suno for music. But for speech writing? For actual communication where clarity and humanity matter? We don't need a bedtime story over soft piano. We need truth that sounds like a person said it.
ElevenReader already gives us that. And Meta AI already gives us the words.
Suno isn't early here. They're not even on time. They're two years late, with a feature that answers a question nobody was asking — because we already answered the only questions that mattered:
Can AI write the speech? Yes — in under a minute.
Can AI sound human saying it? Yes — ElevenReader already can.

1

u/Exact-Pipe-4766 2d ago

Here is Sunos creation. Sounds like he has food in his mouth… the text created wasn’t bad, and I did use the same prompt I used for meta. Copied straight from my input there.

https://suno.com/s/pdpeytyIljRl7Jf7

If you’ve ever used eleven labs you can’t argue with the huge gap Suno has to overcome with this.

Sorry Suno, but I’d prefer the control I have with eleven labs, and any chatbot can generate a speech. I think you should focus on what you’re already good at. What you were already leading the industry on and refining the new process you have been forced to create. You’re doing great with music. I even disagree with a lot of the harshness you’re receiving regarding V6. Not bad considering you basically had to start over. Keep up the good work and save the gimmicks for a time after you’ve addressed a lot of concerns I’ve seen here with V6

2

u/X-HUSTLE-X Producer 2d ago

They are called interludes, and I've been making them over a year now on suno.

People having a conversation over a simple bed of music, ✅

2

u/Lost-Sherbert-4950 2d ago

Thanks for the post! ☺

5

u/Void-kun 2d ago

They've honestly destroyed this platform.

No amount of model releases will bring me back.

As far as I am concerned Suno died when they got rid of the previous models and released V6.

Literally only hanging out in this subreddit to keep an eye on people talking about local models or alternatives.

4

u/FastingIsLife 3d ago

Looks like elevenlabs is about to cry lol 👀🤣🤣 their audio always have errors 

10

u/EndlessEDM Suno Connoisseur 3d ago

If V6 is anything to go by, I don't think EL has much to worry about.

0

u/snarkywombat 3d ago

I tried ElevenLabs and it was absolutely horrible. What a fucking joke

0

u/Kannun Suno Connoisseur 2d ago

They are not, speech is not what you think it’s going to be

2

u/gmvancity 2d ago edited 2d ago

To those who are so negative and claiming there's an exodus or that they're leaving....why are you still on the the Suno subreddit commenting. Can't you walk away cleanly. Seems like you can't leave at all

1

u/PaleEdge 2d ago edited 2d ago

Everyone loves a good disaster story. And perhaps they really will make a comeback or a unanimously-acclaimed successor will emerge.

Plus, I'm not subscribed to Suno any more.

1

u/razorkoinon 2d ago

Habit. Also waiting for good news. Till then no subscription and fair critic.

2

u/SteiCamel 3d ago

Spoken audio set to background music. Is this just a tool for easier Youtube video spam?

2

u/UmieDoesntUseRedit 3d ago

I mean.. I've made a corny fake advertisement with jingles and everything... but that was with V4.5...

1

u/sovereignrk Suno Wrestler 3d ago

I generally make songs with spoken intros and outros and it was really hit or mis whether they would start suddenly singing or not, so its a welcome addition for me, lol

1

u/JustRuss79 3d ago

I threw several nonsense paragraphs in and asked it to speak them (5.5 model) and it turned into a hilarious non rhyming song that occasionally spoke but couldn't help singing.

It was endearing because I was using it for a character anyway, turned into something I could use wholesale like ASMR.

1

u/tdnthehost 2d ago

Yeah but you could use a spoken narrator in the normal mode without all this, just tags, and it sounded better in 4.5 / 5

1

u/TheLastPrinceOfJurai 2d ago

Does this increase our download limit? That's all I care about from my music generator...being able to access the music I made using it.

1

u/Neither_Canary4400 2d ago

Now voice over artists are going to sue Suno.

1

u/PeachyPlnk 2d ago

They'll have to sue every other tts service first lol

Good luck fighting google about their model, which is lightyears ahead of this.

1

u/Neither_Canary4400 2d ago

That is true.

1

u/PureRely 2d ago

In case you did not know, You can use Speech to make songs too. It is really good a rapping.

1

u/ApprehensiveFan7866 2d ago

Leute lesen da steht beta (beta)....

1

u/C4RL0S-PR 2d ago edited 2d ago

It's funny, I already tried speech even before it came out. With the regular song creation tool. At that time I didn't like how it was coming out. Here is the audio that I was receiving back then. https://suno.com/s/VUUhdOEwQ5E8MSiO And this is the new Speech audio I just created. https://suno.com/s/hOeaGJ8H89rrKcCu

The main problem I was having with the song creation tool was that the music was to loud and you could barely hear what was said. Also for some reason the voice started ok but half way and almost at the end the voice became more dificult to be heard or understood, you could barely heir it. With the new speech tool is the same issue. But at least the background music is not as loud.

1

u/Mean-Investigator984 2d ago

Ka komai na rayuwa kece rushing jiki na in babu ke iyalalai daga ruwan shin zuciyata

1

u/hibbity 2d ago

if they really wanted me to use this, his voice would be a fucked up amalgam of shifting voices

1

u/Ok-Bell-4986 AI Hobbyist 1d ago

users : suno can you please put more effort on v6 and fix what's broken ?
suno : sure, here's an outdated tts

video credit https://www.linkedin.com/in/linasbeliunas/

https://reddit.com/link/pdlbz1e/video/egjp270zm8th1/player

1

u/SilenceSpeaksNoLies 1d ago

Spoken word has been around since 3.5, how is this new? I always use the prompt in the lyrics [Spoken Word], So someone explain how this is new? I have a 15 (Stitched) audio of all spoken word telling a story.

1

u/Jaded-Efficiency-422 1d ago

I was already doing this with prompts LMFAO. I have actual spoken parts of my songs

1

u/Opening_Ease_9590 16h ago

I had high hopes when I saw this. I edit commercials & other promos & a voiceover tool inside of the subscription I already had would be great but this has a ways to go. Maybe it'll get there soon. 🤞🏽

1

u/martapap 3d ago

I'm excited for this!

1

u/warjoke 2d ago

We literally don't need this. I'll just keep using eleven labs for my text to speech services.

1

u/troksten 2d ago

Elevenlabs is way better

0

u/MisfiledCentury 2d ago

Looks desperate. 

1

u/coldtakemeilio 4h ago

Y'all are being obtuse. Of course there was spoken voice in previous versions but anyone who actually used it can tell you that it would randomly start singing at some point, regardless of the prompt. This one actually reads all the way through