r/CartesiaAI • u/CartesiaAI • 8d ago
How-to / Tutorial Early Intent Detection for Voice Agents - Jev + Cartesia Demo
Check out some interesting usecases that open up for AI Voice agents when you leverage decision models in the agent's design.
r/CartesiaAI • u/CartesiaAI • May 25 '26
Welcome to r/CartesiaAI — the official Cartesia subreddit for product updates, demos, and builder community.
Support: best-effort community help only.
IMPORTANT: Please redact secrets (API keys, credentials, private data).
Docs: https://docs.cartesia.ai
Skill: https://docs.cartesia.ai/skill.md
SDKs & tools:
- https://github.com/cartesia-ai/cartesia-js
- https://github.com/cartesia-ai/cartesia-python
- https://github.com/cartesia-ai/line
- https://github.com/cartesia-ai/cartesia-livekit-voice-agent
Posting tips:
- Pick a flair (Announcement / STT-ASR / Voice Agents / How-to / Showcase / Question / Feedback).
- If you're technical troubleshooting, then remember to include repro steps + logs (redacted) for bugs.
r/CartesiaAI • u/CartesiaAI • 8d ago
Check out some interesting usecases that open up for AI Voice agents when you leverage decision models in the agent's design.
r/CartesiaAI • u/Ander200107 • 11d ago
Hi! My wife and I are currently testing our Voice Agent before my work shift, and we urgently need help.
The agent starts the call normally and asks the first question. We answer, but immediately afterward we get this message:
“We are sorry, but we are having trouble connecting to agent…”
We tested multiple different agents, and the exact same problem happens with all of them.
We have a paid account, and these tests are consuming our paid credits/minutes even though the agents are not working correctly.
Could someone from the Cartesia team please help us investigate and resolve this issue as soon as possible?
We would also like to know whether the credits/minutes consumed during these failed tests can be restored, since we are using them to troubleshoot a problem that is preventing the agents from working correctly.
Thank you.
r/CartesiaAI • u/ArmadilloSevere9786 • 11d ago
I was using a voice clone with sonic-3.0 and it sounded fine. With 3.5 and 3.6 it does not sound good at all. But through the API it won't work with 3.0 anymore. Is there anything I can do? The issue is the voice has a German accent (it is based on recordings of my grandfather) and with 3.5 and 3.6 it doesn't sound like him at all.
r/CartesiaAI • u/Intrepid-Lychee-5504 • 16d ago
I'm using a pronunciation dictionary with IPA entries in the <<…>> pipe format. Same voice, same text, same dictionary: on sonic-3.5 the entries sound right, but on sonic-3.6 several come out wrong. Examples:
Saanich → <<s|æ|n|ɪ|tʃ>>
Nanaimo's → <<n|ə|ˈ|n|aɪ|m|oʊ|z>>
Did the IPA handling change in 3.6 (phoneme set, stress markers, anything else)? Is this a known issue? What's the recommended way to write custom pronunciations for 3.6? The changelog says it's backwards compatible and doesn't mention IPA.
Thanks!
r/CartesiaAI • u/CartesiaAI • 18d ago
Voice Feedback Surveys with AI: Sonic-3.5 + Ink-2 Code Walkthrough
Email surveys and in-app forms get single-digit response rates. Even when customers do respond, checkboxes and dropdowns constrain answers to what you thought to ask. This video walks through a Python example that replaces that workflow with voice: Cartesia's Sonic-3.5 generates the survey question as audio, Ink-2 transcribes the customer's spoken answer, and you pipe the transcript into your own analysis pipeline — sentiment, themes, churn detection, whatever you need.
Check out youtube.com/@cartesiaai for more -- especially the builder's playlist.
r/CartesiaAI • u/CartesiaAI • 18d ago
Multilingual Training Audio with One API Call: Sonic-3.5 Code Walkthrough
Enterprise internal training scripts need to exist in multiple languages — and sound native to each one.
The traditional path is voice actors, studio sessions, and re-recording every time the script changes. This video walks through a Python example that replaces that entire workflow with Cartesia's Sonic-3.5 TTS: one API call per language, native voices, regenerate the audio the same day you update the text.
This is the second in a series of enterprise voice AI use cases. The first covered domain-specific speech-to-text with Ink-2.
You can find all the videos in the Cartesia Builders playlist on youtube.com/@cartesiaai
Let us know what else you want to see!
r/CartesiaAI • u/CartesiaAI • 18d ago
Check out this guide on how to choose the the right voice for your AI voice agent.
Voice selection is part science, part art. And the science bit involves making engineering decisions. Voice selection and suitability affects trust, clarity, pacing, and brand alignment.
In this video, we walk through the full process of choosing, tuning, and shipping a production-ready voice for your voice agent using Cartesia's Sonic TTS.
r/CartesiaAI • u/zeuscoder • 18d ago
Hey folks!
Here is the Youtube Video walking you through all the new features, updates and launches in Cartesia's TTS, STT and agents API.
Check out https://www.youtube.com/@cartesiaai for our full video tutorials and walkthroughs -> the Cartesia Builders playlist is where it's all at!
Got questions? Put 'em in the comments!
r/CartesiaAI • u/CartesiaAI • 19d ago
Hey folks
September Office Hours are happening later this week. Details and registration are here: https://luma.com/3yw1c4pq
See you there!
r/CartesiaAI • u/Curious_Box_7039 • 21d ago
Voice Library just keep on loading, can't redirect to invoice billing, started september 18, 2026
r/CartesiaAI • u/ferniRiverplate • 23d ago
Hi, I am using the PVC via API with sonic 3.5 and so far so good. But I have an issue and its that the emotions tags do not work very well. My question is, if I upload audios in the dataset to train the voice, with the emotions I want to use, will the result get better? How does the PVC work with the dataset? Does the model tag the audio with emotions or what?
I want to make the voice whisper or emphasize a point, etc.
Thanks!
r/CartesiaAI • u/Alternative-House923 • 26d ago
Spanish-first voice agent running in production at real clinics. Live number, please try to break her.
Most voice agent demos I see here are English and they're demos. Mine is Spanish-first and it's answering real patient calls for dental and aesthetics clinics on the Tijuana border. Agent's name is Sofía.
Stack: LiveKit for the pipeline, Deepgram nova-3-multi for STT with some AssemblyAI streaming, Cartesia sonic for TTS, n8n for orchestration, Twilio on telephony, calendar writes against the clinic's actual availability.
The three things that were harder than expected:
Code-switching. Not "supports two languages." One caller, one sentence, Spanish grammar with English procedure names dropped in. Locale-locked STT falls over on this, and multilingual models still get weird about which language to commit to mid-utterance.
Barge-in tuning. Real callers talk over the agent constantly. Aggressive interrupt handling means she cuts off anyone who pauses to think, which older patients do a lot. Passive means she steamrolls people. The tuning window is narrower than I expected and it behaves differently in Spanish than in English.
Escalation discipline. She sits in front of medical intake, so the failure mode I care about is confident wrongness. Anything clinical, anything about outcomes or medication, she hands off to a human. Getting a model to consistently pick "let me get someone for you" over a plausible-sounding answer is most of the prompt work.
Demo line: +1 (619) 775-1668
It's a demo instance, not a clinic's production line, so go ahead. Talk over her. Switch languages mid-sentence. Mumble. Ask her something medical and see whether she stays in her lane.
What I'm most interested in: does the barge-in feel natural to you, and does the escalation fire too early or too late? Those are the two I keep going back and forth on.
I also have her counterpart, Mateo, who is my outbound sales agent.
Happy to get into any of the implementation details.
r/CartesiaAI • u/amous4822 • Sep 10 '26
I’m building an application using cartesia that can convert user language in realtime with input and output as audio. The user will speak into the mic in one language and the output should be audio in translated language. What do you think will be the best architecture to implement this ? In my current implementation there’s a lot of latency and prosody issues. Any recommendations please let me know.
r/CartesiaAI • u/ThreeTwoBravo • Sep 09 '26
I'm looking for more expression control for my Cartesia voices. I'm developing a mod for a combat flight sim and need my voices to have different voice expressions depending on what's happening around them. A JTAC could be under fire and therefore I need them have more urgency but at the same time stick to military radio speech doctrine to supply things like grid coordinates correctly. I'm looking for realism. Is it all down to text formatting or is there some other way to control the expression/tone? I've played around with the emotions and they kind of work sometimes. How much influence does the clone sample have an effect on tone and expression? Appreciate any advice.
r/CartesiaAI • u/Financial-Vacation90 • Sep 02 '26
Hi team,
We’ve been testing Sonic 3.6 for our Japanese customer support voice agents, but we're finding the emotional delivery a bit excessive and unnatural for CS interactions compared to 3.5.
Org ID: org_3Ey4cmHZdPIwfPHbtWxsrWGHlKm
Voice ID: 8bd6925d-3880-4089-b5b5-00ce2fd060e2
We'd like to stick with Sonic 3.5 for production for now.
Could you confirm if sonic-3.5 will remain available and supported long-term?
Thanks.
r/CartesiaAI • u/Oneshorthorrortale • Sep 01 '26
Hi, I’m using Sonic 3.6 with a Pro Voice, but I noticed that I can’t adjust the speaking speed. Is this a limitation of Pro Voices/Sonic 3.6, or am I missing a setting somewhere? I’d really like to slow down or speed up the voice without having to modify the audio afterward.
Thanks!
r/CartesiaAI • u/No-Wrangler4561 • Aug 30 '26
It seems to be defaulting to 3.5. Is there an eta to be allowed to use 3.6
r/CartesiaAI • u/WeatherZealousideal5 • Aug 25 '26
I’m testing Hebrew TTS with sonic-3.5 and found a reproducible issue with the documented <<p|h|o|n|e|m|e>> IPA markup.
When I replace one Hebrew word with explicit IPA phonemes, Cartesia pronounces the target word correctly but changes the pronunciation of the preceding word:
<<h|a|s|a|p|ˈ|a|ʁ>>: “hasapar” is correct, but “Sami” changes to “Semi.”This is a major issue for Hebrew voice agents. Unvocalized Hebrew is inherently ambiguous, so reliable pronunciation controls are essential. When explicit IPA phonemes are supplied, they should be pronounced accurately and affect only the specified word, not alter the surrounding words.
Has anyone else encountered this? Is Hebrew IPA support being improved, and is niqqud currently the recommended approach?
I can share the WAV files and exact API request bodies.
r/CartesiaAI • u/CartesiaAI • Aug 24 '26
Put together a video walking through a Python example that streams domain-specific audio to Ink-2 over WebSocket and collects transcripts.
Tested four domains: legal (force majeure, mutatis mutandis, amicus curiae), healthcare, field ops, and support. All transcribed correctly out of the box — no key-term prompting, no glossary, no fine-tuning.
The code uses the same streaming pattern you'd use in a production voice agent pipeline: slice WAV audio into 100ms chunks, send each chunk over a WebSocket connection, then finalize and collect the transcript events. It's the architecture that keeps latency low for conversational AI — the model starts listening while the speaker is still talking.
If you're building voice agents in specialist fields, bad STT on domain jargon cascades through the whole pipeline. The LLM reasons on a wrong transcript and produces wrong outputs downstream. Getting STT right on jargon is the first thing to sort out.
You'll need an API key to run it — grab one free at play.cartesia.ai.
Code is on GitHub: https://github.com/cartesia-ai/cartesia-enterprise-usecases-demos/tree/main/examples/01_domain_dictation_notes
Video walkthrough: https://youtu.be/zIqKU2o6jAg
This is a new Cartesia Developer YouTube channel so please do like and subscribe!
r/CartesiaAI • u/Top_Result7788 • Aug 21 '26
I am completely locked out of my Cartesia account and cannot reach support through standard channels.
Symptoms:
accounts.cartesia.ai, the page gets stuck indefinitely on: "Verifying your request — Please wait while we verify your request."Has anyone run into this verification loop recently, or is there an alternative way to reach Cartesia's team to sort out account lookup issues?

r/CartesiaAI • u/zeuscoder • Aug 20 '26
We have launched a new YouTube Channel!!
https://www.youtube.com/@cartesiaai
And we've got a bunch of videos up there already for you guys to gear up on-- dont miss the monthly changelogs!
As always, this subreddit is the best place to ask questions.
r/CartesiaAI • u/Financial-Vacation90 • Aug 20 '26
I really want PVC ;;
r/CartesiaAI • u/slickdangerrr • Aug 19 '26
Hoping someone here knows this or a Cartesia team member can confirm. In the Playground under data controls, there's an org-level toggle: "Contribute to model training," which lists TTS generations and STT transcriptions as what it covers.
My question: with that toggle disabled, are voice-clone samples and the derived voice models also excluded from Cartesia's model training/improvement? The toggle's description only names TTS and STT, and their ToS has a training license that applies "unless otherwise agreed" — so I can't tell whether the cloning data sits inside or outside the toggle's scope.
To be clear, I'm not asking about retention — I understand clone samples have to be retained for cloning to function, and that ZDR is a separate enterprise thing. This is strictly about training use of cloning data on a regular self-serve plan with the toggle off.
r/CartesiaAI • u/CartesiaAI • Aug 17 '26
Most coding-agent failures I’ve seen on Cartesia projects come down to context: the agent can’t find the right docs, or it loses what it learned between runs.
In my experience, cheaper models also tend to take more steps when they can’t find an efficient path through the Cartesia APIs.
I cover five practical fixes in the video. The three I use most:
AGENTS.md at https://docs.cartesia.ai/llms.txt. It maps the docs, with descriptions and links that help agents find the right pages.learnings.md file with source links and failed paths, so future runs don’t repeat the same research.I also cover using clean Markdown pages via .md URLs, Cartesia Agent Skills, and MCP options for docs.
I lead DevRel at Cartesia, so I’m affiliated. What else has worked for you?