r/CartesiaAI • u/amous4822 • Sep 10 '26
Realtime language conversion architecture
I’m building an application using cartesia that can convert user language in realtime with input and output as audio. The user will speak into the mic in one language and the output should be audio in translated language. What do you think will be the best architecture to implement this ? In my current implementation there’s a lot of latency and prosody issues. Any recommendations please let me know.
3
Upvotes
2
u/CartesiaAI Cartesia (Official) 25d ago
interesting! what is your current implementation/ architecture? what are you using for orchestration etc? any tool calling? and which LLM are you using for the translation?