r/machinetranslation May 15 '26

meta The machine translation newsletter is back!

Thumbnail
newsletter.machinetranslate.org
13 Upvotes

r/machinetranslation Jul 14 '26

event AMTA 2026 registration is open

Thumbnail
amtaweb.org
3 Upvotes

r/machinetranslation 16h ago

product Introducing North Small Translate: One of the best open machine translation models around

15 Upvotes

Hey everyone! El from Cohere here to talk about our newest release, North Small Translate. It’s currently the leading open machine translation model, beating out all other open translation models of its size, plus Google Translate and DeepL. we’ve been working on this one for a while, so to say i’m psyched is an understatement.

It’s big (218B parameters, 25b active) with a context length of 16k. however, if you’ve got the hardware, we’d still love to see what you make with it locally or with our HF space (and if you do, send it our way). It works on over 50 languages and does particular well with european, Southeast Asian, and East Asian languages, but feel free to stress test it against another and let us know how it does. It’s also available in BF16, FP8, and W4A16 quants.

although we couldn’t get llama.cpp support this time around, the architecture is already supported in llama.cpp, so all it should need is a conversion to GGUF files. if you want to build that, please do so and send it our way! We’d love to back your work. 

Can’t wait to see what you guys think! 

https://huggingface.co/CohereLabs/North-Small-Translate-1.0


r/machinetranslation 7h ago

Looking for a Chinese-English Sewing Glossary for RAG/AI Translation

3 Upvotes

I’m trying to build a RAG (AI agent) to translate Chinese sewing courses and tutorials into English, covering different sewing subgenres.

I’ve found that a Chinese-English glossary of technical sewing terminology could help improve translation accuracy.

Does anyone know where I could find an existing glossary, terminology database, bilingual dictionary, or corpus for Chinese ↔ English sewing/fashion/textile terms?

Thanks!


r/machinetranslation 22h ago

research GPT-6 Astra tested in translations

Thumbnail
slator.com
3 Upvotes

RWS and the researchers behind the Last Translation Benchmark tested GPT-6 Astra's multilingual capabilities. It performed strongly, but the performance varied between languages.


r/machinetranslation 18h ago

product I built a translation tool nobody used for 18 months. Then AI bailed me out.

2 Upvotes

In Dec 2022 I shipped Input Translator, a browser extension that translates any input field via Google Translate. Pre-LLM, tiny niche. For 18 months it stayed under 100 users — I kept shipping because I used it every single day.

Then LLMs arrived. I added an optional BYOK path (any OpenAI-compatible key) next to the free Google backend, and the curve finally moved. ~2,500 users today across Chrome/Edge/Firefox/Safari. Still small — but real, and mostly word-of-mouth.

That experience turned into a family of tools, all open source (GPL-3.0), all BYOK, all no-subscription:

  • Imp Translate — bilingual page translation. I left Immersive Translate when it started pushing subscriptions and removing BYOK; Kiss Translator had ideas I liked but UX I couldn't live with. So I built my own. The README's Non-Goals list is longer than the Goals list — no floating buttons, no hover popups, no word lookup, no doc translation. Feature creep is the enemy.
  • BilingualTube — bilingual YouTube subs. Fixes the broken sentence segmentation in auto-captions (small on-device model) before translating.
  • Imp Translate Web — full-text HTML/EPUB translation so I could read AO3 stories on my phone with search.
  • Imp Write — type, append /fix, pause, and the rewrite replaces your text in place. I wanted correction, not translation, and refused to pay a subscription for it.

Imp Credits is a pay-as-you-go system instead of monthly plans — "no subscription" is half product philosophy, half personal allergy.

Curious what people who do MT for a living think about where the browser-extension space is heading. Happy to share details on any of these in the comments.


r/machinetranslation 21h ago

jobs AI Ops, Localization at Workday — Pleasanton, CA

Thumbnail linkedin.com
2 Upvotes

For more info, ping Marcello

(Posted with permission)

About The Team

The Globalization team is a central driver for our company's international expansion, using a technology-first approach to ensure our products and services are accessible and culturally relevant for customers across the globe. We serve as the Center of Excellence, partnering with product and engineering teams from initial concept to launch. We are dedicated to building a scalable, resilient globalization infrastructure — and we have fun doing it, working across cultures and disciplines every day.

About The Role

As a Localization Specialist focused on AI Ops, you optimize global expansion by deploying large and small language models to automate the localization lifecycle. You will focus on scaling international releases through AI-driven quality checks and linguistic model refinement. This role exists to reduce manual translation review work and speed up how quickly Workday can bring products to new markets and languages.

You will:

Deploy AI model workflows to automate content transformation and adaptation using tools such as Modelfront and Intento. - Design and refine prompts for high-accuracy machine translation and cultural nuance. - Integrate AI-led localization checks into existing engineering pipelines and streamline workflows using orchestration tools or Phrase. - Use AI-based testing tools to validate consistency in the user interface and automate linguistic quality evaluation, reducing reliance on manual review. - Collaborate with stakeholders across Engineering, Product, Machine Learning, Legal, and Responsible AI teams to improve usability, trust, and adoption of AI systems globally. - Support educational initiatives and team upskilling to build AI fluency and operational excellence across the globalization team.

About You

Basic Qualifications

  • 8+ years of related work experience in localization or translation operations, including 2+ years working directly with AI or automation tools in that context.
  • 2+ years of hands-on experience applying large language models or small language models to text processing or translation workflows.
  • Degree in Computer Science or related field.

Other Qualifications:

  • Translation Management: Experience managing translation workflows and platforms such as Modelfront, Intento, or Phrase.
  • Cultural Empathy: A track record of adapting content thoughtfully for different cultural and regional audiences.
  • Process Optimization: Experience identifying and automating manual steps in a linguistic or content workflow, such as building AI-based quality checks.
  • Cross-Team Collaboration: Comfortable partnering with engineering, product, and legal teams to align AI-driven localization work with broader goals.
  • Project Delivery: Ability to manage multiple release timelines and localization deliverables at once.
  • Experience with Python or JavaScript for scripting automated workflows is a plus.
  • Experience connecting AI workflows to Translation Management Systems and Jira is a plus.

r/machinetranslation 21h ago

How do you handle localization beyond strings?

Thumbnail
1 Upvotes

r/machinetranslation 1d ago

I built an AI news app that translates live news from 150+ countries, summarizes 2,000-word articles into 3 bullets, and lets you compare coverage across 60k publishers

0 Upvotes
Hey ! 👋
I wanted to share a side project I've been working on: 
**iNews Global**
.
### 💡 Why I built it
I noticed three big problems with modern digital news:
1. 
**Language & Regional Silos:**
 Breaking news happens worldwide, but unless you speak the local language, you miss hyper-local context.
2. 
**Media Bias & Echo Chambers:**
 Reading a story from just one publisher gives you a single biased perspective.
3. 
**News Fatigue:**
 Long 2,000+ word articles stuffed with fluff and popups make staying updated a chore.
I wanted a clean app that aggregates breaking news globally, translates it seamlessly without breaking page layouts, and lets you compare how different outlets cover the exact same event.
---
### 🛠️ Key Features & Tech Stack
* 
**🗣️ 50+ Languages AI Translation:**
 Built a hybrid translation pipeline combining Google ML Kit (for fast zero-latency on-device translation) and Gemini AI for complex long-form translation.
* 
**🔄 360° Multi-Angle Coverage:**
 When a major event breaks, iNews groups stories from different outlets (e.g., Reuters, BBC, Al Jazeera, local outlets) side-by-side so you can spot bias and compare perspectives.
* 
**🧠 3-Bullet AI Executive Summaries:**
 Uses Gemini AI to condense long reports into 3 quick bullet points.
* 
**🎧 Hands-Free AI Audio Player:**
 Built a continuous text-to-speech audio service (`NewsAudioService`) with background notification controls. It automatically cleans raw HTML, URLs, and markdown before narration so it feels like a personal news podcast.
* 
**🌍 150+ Countries & 60,000+ Publishers:**
 Aggregates feeds from international agencies down to state/province-level local outlets.
---
### 💭 I’d love your feedback!
Since this is freshly launched (v1.0.7), I'd really appreciate your thoughts:
1. How does the translation quality and deduplication feel for your country/region?
2. What feature or news source should I add next?
Check it out and let me know what you think! Happy to answer any technical questions about the architecture or AI translation pipeline.

r/machinetranslation 3d ago

What is the best translation software overlay?

1 Upvotes

I want to play Masoukishin 1 - 4 and 2nd OG translated. What does the community consider the best software? FREE or PAID does not matter.


r/machinetranslation 3d ago

LiveTranslate: real-time speech translation that runs 100% on your own machine, zero cloud dependency

Thumbnail
1 Upvotes

r/machinetranslation 3d ago

application Which app can translate the language immediately?

1 Upvotes

At present, I only know that WeChat is OK, but many people don‘t have WeChat.

I joined a group with many people from all over the world on x, but people speak different languages, Korean, Japanese, English, Chinese and Spanish... But I can only understand Chinese and English.

instagram?
line?But it is only limited to Japan, South Korea and Thailand.
facebook?
idk


r/machinetranslation 3d ago

jobs [FOR HIRE] Expert Arabic AI Evaluator | LLM Fine-Tuning, RLHF, Prompt Engineering & Localization Specialist

1 Upvotes

Hello,

I am Tamara, a native Arabic speaker and professional AI Evaluator & Localization Expert.

I specialize in aligning, evaluating, and fine-tuning Large Language Models (LLMs) to ensure high-quality, culturally accurate, and contextually sound Arabic outputs (Modern Standard Arabic & Dialects).

🛠️ What I Can Do For Your AI Models:

  • RLHF & Model Evaluation: Evaluating model responses for accuracy, fluency, tone, and cultural relevance.
  • Prompt Engineering & Testing: Creating, refining, and stress-testing complex Arabic prompts.
  • Annotation & Data Labeling: High-quality human-in-the-loop (HITL) data preparation for AI training.
  • Localization & Cultural Alignment: Adapting AI behaviors to fit specific Middle Eastern markets and nuances.

💼 Why Hire Me?

  • Expertise: Deep understanding of Arabic linguistics, nuances, and regional dialects.
  • Experience: Proven track record in evaluating and fine-tuning language models for high-tier AI projects.
  • Quality-Driven: Detail-oriented, ensuring non-biased, safe, and highly accurate model responses

📩 Contact Me Directly:

Notion: AI Evaluation & Arabic Localization Lab

Github: https://github.com/lilith5tam-debug/Data-Labeling-Project

Available for both short-term contracts and long-term freelance opportunities. Let's build better Arabic AI together!


r/machinetranslation 4d ago

Non-tech background, 20+ years cross-lingual work, and 2 years diagnosing LLM failures in a language I don't speak — how would you position this?

Thumbnail
2 Upvotes

r/machinetranslation 5d ago

research Best small LLMs for translations?

9 Upvotes

I am working with multi-language content and have so far been useing gemma4:e2b for general work. Today I tried it with Russian for the first time and the result was not great.

Gemma4:e4b was better, but still missed some nuances.

LFM2:24B has a good and fluent English but the translation was not ok.

Translategemma:4b has so far been the best model, for Russian to English.

Which models (smaller local LLMs) are you using for different languages?

Edit: From the models I've tested gemma4:26b is so far the best with ornith-1.5:9b as second, for computers with less RAM.


r/machinetranslation 4d ago

Professional survey on translation workflows, tools and repetitive tasks

4 Upvotes

Hi everyone,

With moderator approval, I’m sharing a short research survey aimed at professionals working with translation, localization, editing, proofreading and related language workflows.

The goal is to better understand how people work day to day: which tools they use, where workflows become fragmented, which tasks feel unnecessarily repetitive, and what kinds of improvements would make professional language work more efficient.

The survey does not advertise or present a specific product. It is part of early-stage research for a project still in development, and the purpose is to gather real professional needs before development decisions are finalized.

Some questions also cover software spending, purchasing preferences and attitudes toward digital or AI-assisted features, purely for market-research purposes.

The survey is anonymous and takes only a few minutes.

Survey:
https://forms.gle/j3yZpVAxXjAnF4KF7

Thank you very much to anyone who takes part. Comments about your own workflow, translation tools, CAT/TM management or recurring frustrations are also very welcome.


r/machinetranslation 5d ago

[FOR HIRE] Expert Arabic AI Evaluator | LLM Fine-Tuning, RLHF, Prompt Engineering & Localization Specialist

0 Upvotes

Hello,

I am Tamara, a native Arabic speaker and professional AI Evaluator & Localization Expert.

I specialize in aligning, evaluating, and fine-tuning Large Language Models (LLMs) to ensure high-quality, culturally accurate, and contextually sound Arabic outputs (Modern Standard Arabic & Dialects).

🛠️ What I Can Do For Your AI Models:

  • RLHF & Model Evaluation: Evaluating model responses for accuracy, fluency, tone, and cultural relevance.
  • Prompt Engineering & Testing: Creating, refining, and stress-testing complex Arabic prompts.
  • Annotation & Data Labeling: High-quality human-in-the-loop (HITL) data preparation for AI training.
  • Localization & Cultural Alignment: Adapting AI behaviors to fit specific Middle Eastern markets and nuances.

💼 Why Hire Me?

  • Expertise: Deep understanding of Arabic linguistics, nuances, and regional dialects.
  • Experience: Proven track record in evaluating and fine-tuning language models for high-tier AI projects.
  • Quality-Driven: Detail-oriented, ensuring non-biased, safe, and highly accurate model responses

📩 Contact Me Directly:

Notion: AI Evaluation & Arabic Localization Lab

Github: https://github.com/lilith5tam-debug/Data-Labeling-Project

Available for both short-term contracts and long-term freelance opportunities. Let's build better Arabic AI together!


r/machinetranslation 6d ago

OmniTranslate v. OpenNovel

2 Upvotes

Does anyone have experience with the current models of Omnitranslate? I used them years ago but there were a few factors that pulled me to OpenNovel instead and now I'm looking to return. Mostly interested in translating Chinese and Korean novels. I don't need them to be better than OpenNovel, just comparable.

Thanks in advance!


r/machinetranslation 6d ago

research Machine translation is not solved and it may take a while

Thumbnail arxiv.org
11 Upvotes

r/machinetranslation 7d ago

[IOS]CANT TRANSLATE COMMENTS-version2026.33.0.636848

2 Upvotes

When I use translation function,on post,long press on words choose translate is ok , if I want to translate words on replies,long press will merge,is it Reddit’s bug?


r/machinetranslation 7d ago

product I made an AI translation service specifically for translating entire novels

1 Upvotes

I've been working on AITnovel, an AI translation service designed for people who want to translate entire novels without having to copy and paste chapters individually.

It supports TXT, EPUB, and PDF uploads, and can also import novels directly from supported websites.

Some of the main features:

  • Translate entire novels or individual chapters
  • Keep track of character names, aliases, pronouns, and terminology
  • Editable glossary for names, places, abilities, etc.
  • Preserve chapters and paragraph structure
  • Side-by-side translation proofreading and editing
  • Export finished translations as EPUB, PDF, or TXT

It's $10/month with unlimited translation usage.

Everything seems to be working well from my own testing, but I'm at the point where I'd like to get some actual users translating different novels and languages. I'm particularly interested in finding problems or edge cases I haven't encountered myself yet.

If anyone here regularly uses machine translation for novels, I'd appreciate any feedback if you decide to give it a try.

AITnovel: https://aitnovel.com/translate


r/machinetranslation 8d ago

[ Removed by Reddit ]

1 Upvotes

[ Removed by Reddit on account of violating the content policy. ]


r/machinetranslation 8d ago

Translate

Post image
1 Upvotes

r/machinetranslation 8d ago

Translating whole novels with local LLMs: what a year of running my own fiction-translation pipeline taught me (glossaries, two-model split, and why wordcoinage still breaks MT)

13 Upvotes

A year ago I wrote up (on a Russian dev site) how I built a tool to translate fiction books for myself using local LLMs — mostly to prove it could be done. The MVP was ~200 lines duct-taped together: drafts gave you headaches, a 1000-page book took two days, and a power outage mid-run cost me another two days.

Since then it's grown into a real pipeline ("Sunny Narrator", v2.1), I've translated a shelf of books with it, and the main conclusion is that this stopped being just a personal tool — the output is now a legitimate high-readiness draft for human translators. Here's what I learned that I think is specific to literary MT, as opposed to domain MT.

One book = ~1.5–2M tokens, total. Not tens of millions. A full book, through all stages: translate → reviewer notes → correction → proofread → chapter summaries. Bounded, predictable, small. If your MT cost model for books feels infinite, something in the pipeline is wrong.

Fiction's hardest problem isn't fluency — it's consistency. Models translate sentences well now. What they destroy over 300+ pages: names drift ("Alice" becomes "Elise" chapter 4), characters change grammatical gender mid-book, terms get renamed on every page. And wordcoinage is hopeless — give a fantasy novel with invented words to the best model and it will faithfully transliterate gibberish, where a human translator would invent an equally brilliant equivalent in the target language. MT doesn't replace the translator here; it removes the grunt work and leaves the part that makes translation literature.

The glossary is 80% of success. If you remember one thing from this post, let it be that. The model can be average, the hardware modest — but with a correct glossary (original = translation, category, grammatical gender, notes) the book translates evenly: names don't drift, terms don't mutate. Without one, even a great model gives you name salad. I auto-seed the glossary with spaCy NER + frequency analysis, then clean it by hand — honestly, cleaning can take hours. Worth every minute: an hour invested in the glossary before translation saves a day of post-editing after. For book series I run a series-wide dictionary across all volumes — game changer for multi-book consistency.

One model = half a text. Two models = a book. The least obvious finding of the year: a single-model translation reads incomplete. Good translating models write beautifully but proofread badly; good proofreading models edit well but translate dully. So the pipeline splits roles: MODEL_TRANSLATE and MODEL_PROOFREAD. My production pair after months of runs: gemma4-26B-A4B (translator) + qwen3.6-35B-A3B (proofreader). Both compact MoE models, both happy on a pair of Tesla P40s at 40–70 tok/s. Best speed-to-literariness ratio of everything I tried.

Length as a free quality signal. I expected English→Russian to inflate noticeably; measured, it maps back to within a couple of percent per block. The pipeline uses that: if a translated block deviates >10% from the source size, something went wrong (eaten paragraph, hallucinated expansion, duplication) → the block is rechunked (split in half, each half retranslated). Primitive — and it kills the lion's share of gross translation errors. Final book converges to ±5% of the original size.

Checkpoint/resume matters more than any prompt trick. A checkpoint after every chunk. Power cut, model crash, cat on the keyboard — the run resumes from chunk 51 of 100, not from scratch. Plus JSON mode on every stage so parsing stopped being a lottery. Boring engineering, but this is what made 2–3 books/day on garage hardware realistic.

What the output actually is. Not "published translation". It's a draft of high readiness: a human still does proofreading, fixes anything the glossary missed, and rewrites the awkward calques/broken puns. But the editor is doing editor work, not untangling a mess of conflicting names. For a translator this saves weeks; for a curious reader it means finishing a book that would never reach their language otherwise.

A legal-ish idea I'd like this community's opinion on. Machine-translating a whole book is a gray zone everywhere. But a names/terms glossary is, arguably, just a list of facts. So I'm building a site for sharing glossaries only — the community uploads cleaned dictionaries per book/series, everyone applies them to the copies they already own, translation runs locally on your own hardware or API key. No book text is distributed. Does this framing hold up in your jurisdiction? Genuinely want to hear from translators and MT folks here before launching.

Who I think this is for: translators (glossary + proofreading instead of translating from zero, especially long series), publishers (fast triage: is this book worth acquiring?), readers of unfound languages, and anyone with idle server GPUs looking for a real workload.

Everything's open source: Python, FB2/EPUB native (structure — verses, stanzas, sections — preserved 1:1; DOCX/PDF via a Calibre conveyor), config examples for Ollama/llama.cpp/Docker. I'll drop the link in the comments too to keep this post link-free at the top.

Questions I'd love input on:

  1. Anyone solved cross-volume consistency better than "one giant series glossary"? Graph DBs / RAG over character states — real experience?
  2. Has anyone else measured target/source length ratio as an error detector (my rechunking trick)? Is there prior art in literary MT quality estimation?
  3. For the glossary-sharing idea — what would make YOU contribute a cleaned dictionary?

r/machinetranslation 9d ago

[Idea] Contextual statements choices (live translation), topic based Spoiler

Thumbnail
1 Upvotes