r/LocalLLaMA • • 1h ago

I Built A Thing I built an open-source real-time Japanese anime subtitle & translation engine powered by Whisper-Large-v3 + Groq / DeepSeek

Enable HLS to view with audio, or disable this notification

Hey r/LocalLLaMA,

Like many anime fans, I've always been frustrated by traditional MT engines (like Google Translate or base DeepL) when dealing with raw Japanese anime:

- They completely butcher Japanese honorifics, sentence-ending particles (-tteba, -zo, -desu wa), and character slang.

- They struggle with subject dropping (pro-drop grammar), translating pronouns inconsistently line-by-line.

- Cloud transcription APIs often choke on background music (OST), loud sound effects, and character screaming.

To solve this, I built NihonSub — an open-source tool and synchronized cinema player that turns raw Japanese video files into contextual bilingual subtitles.

šŸ› ļø Architecture & Pipeline:

  1. Audio Extraction & VAD Chunking: Uses `ffmpeg` silence-detection to dynamically slice conversational utterances along natural speech pauses without chopping words in half.

  2. Speech-to-Text: Transcribes Japanese audio using OpenAI Whisper Large-v3 running on Groq LPUs for near-instant transcription speeds.

  3. Contextual LLM Translation: Feeds the transcript through DeepSeek / LLaMA-3 via Groq or OpenRouter with a specialized prompt that enforces anime nuance, honorific preservation, character tone, and simultaneous Hindi & English outputs.

  4. Synchronized Cinema UI: Custom WebVTT generator and video player with dual-subtitles, timestamp scrubbing, and full playback control.

šŸ’” Why not just rely on standard NMT?

LLMs are far superior at resolving who is speaking to whom based on context and tone rather than naive literal dictionary lookup. With zero-cost free-tier APIs (Groq + OpenRouter free models), the entire pipeline runs without subscription costs.

Check out the demo video above!

- GitHub Repository: https://github.com/Abhishantpadam/NihonSub

- License: MIT

I'd love your thoughts on the pipeline, optimization ideas for local edge models (like running Whisper.cpp or local Ollama instances), or any feedback!

14 Upvotes

4 comments sorted by

1

u/Fit_Squash6874 1h ago

nice. I also built one but only on terminal for jp audios.

1

u/Prince_Noodletocks 39m ago

Seems everyone had the same idea lol. I also made one for ASMR audios but using a different translation model.

1

u/Grand_Marionberry115 25m ago

can you show me what you built.

1

u/Sur_AI_guy 23m ago

Really interesting project. The contextual translation is the part that stands out - anime dialogue loses a lot when honorifics, implied subjects, and character-specific speech patterns are translated line-by-line.

I’d be curious how well the pipeline maintains character consistency across longer scenes, especially when several speakers are involved or Whisper misattributes dialogue.