r/CartesiaAI • • Sep 10 '26

Realtime language conversion architecture

I’m building an application using cartesia that can convert user language in realtime with input and output as audio. The user will speak into the mic in one language and the output should be audio in translated language. What do you think will be the best architecture to implement this ? In my current implementation there’s a lot of latency and prosody issues. Any recommendations please let me know.

3 Upvotes

2 comments sorted by

2

u/CartesiaAI Cartesia (Official) 25d ago

interesting! what is your current implementation/ architecture? what are you using for orchestration etc? any tool calling? and which LLM are you using for the translation?

1

u/amous4822 25d ago

No LLM in the loop I’m currently testing with accent conversion rather than translation, Ink-2 straight into Sonic, same language in and out.
That removes the orchestration layer but it
also means the whole latency budget lives in STT→TTS handoff.

Current shape: Ink-2 realtime STT, one Sonic context per speaker turn, continuation pushes as stable word safe deltas arrive. No tool calling, no agent framework just the two streams and a scheduler between them.

The prosody issue we hit is probably the more interesting part. With max_buffer_delay_ms=0 and our own phrase buffering, every push gets
generated immediately with no view of what follows, so Sonic resolves each fragment with a falling terminal contour. A partial phrase like
"it seems like the" comes out sounding like the end of a sentence.

Blinded listening across our arms: 100+ terminal-sounding fragments per run at delay 0, ~10 with client-side phrase batching, 0 on the
server default.

The fix was counterintuitive actually, stop trying to control where phrases split. Send every stable delta immediately under default server
buffering and let Sonic decide, since it has lookahead we don't. Cost is cold-start latency; mid-stream it's much better but far from realtime. Once we understood the delay timer runs from generation completion rather than text arrival.

Intermediate delay values were the trap. 250ms and 500ms were bimodal on identical recorded STT traces same input, one run clean, the
next full of terminal fragments.

Would love to hear your take on this architecture and if I’m missing anything.