Training Custom MT on Your Translation Memories: A Practical Guide¶
Your agency has 500,000 segments in translation memory. When you feed a new document to Google Translate, the engine has no idea about your memories. It translates fine but misses your client’s tone, specialized vocabulary, and learned nuances.
That’s where custom MT training enters. You train your own engine on your translation memory. The engine learns your terminology, style, and domain expertise. Result: faster post-editing, higher consistency, measurable turnaround gains.
What is Custom MT Training?¶
Custom MT training fine-tunes a neural MT engine on your translation memory. The model learns domain-specific patterns from your data.
A generic engine (Google Translate, DeepL, Microsoft) trains once on billions of public words and never changes. A custom engine is a new model trained specifically on your data. It retains base knowledge but layers your domain expertise on top.
ROI Threshold¶
The ROI threshold is roughly $60,000 in annual translation spending per language pair.
An average training job costs $300-400. Three language pairs = $900-1,200 per initial training. Add infrastructure ($0-2,000/month) and you’re looking at $1,500-3,000 to start.
Post-editing effort typically drops 15-40% with high-quality data. For an agency with $60k annual volume, 20% effort reduction = ~$12,000 recovered annually. Training pays back in 1-3 months.
Bottom line: Under 50,000 words/month is optional. Over 100,000 words/month is usually smart.
Data Requirements¶
Minimum: 6,000 translation units. Recommended: 10,000+ units (~150,000 words).
Why? Neural networks need patterns. Quality matters more than size—dirty data reduces performance.
The Process¶
- Prepare and upload TM (2-4 hours)
- System processes data (30 min-2 hours)
- Training begins (1-1.5 hours for ~1M words)
- Validation and deployment (15-30 minutes)
- Monitor and retrain (every 6-12 months)
Platforms¶
- Google Translate AutoML: $300-400/training, Google Cloud setup
- Microsoft Custom Translator: ~$300/training, Azure setup
- Smartling: Bundled TMS pricing, integrated setup
- Phrase Custom AI: ~$300-500/training, automated curation
- ModernMT: Per-segment model, real-time adaptation
- Custom.MT: $39-45/hour, simplified interface
Common Mistakes¶
- Dirty TM: Remove outdated translations, inconsistencies, low-confidence segments
- Too small dataset: Aim 10,000+ minimum
- No baseline: A/B test before and after
- One-time training: Retrain every 6-12 months
- Mixed domains: Train separate engines per domain
Measuring Success¶
- Post-editing distance should reduce 15-40%
- Keystroke/effort metrics should decrease
- BLEU score should improve 10-25%
- Turnaround should be 20-30% faster
The Real Example¶
Legal agency, 300,000-segment English-German TM:
- Phase 1: Prepare and upload (4-6 hours)
- Phase 2: Training (2-4 hours, $300)
- Phase 3: Integration and testing (2-3 hours)
- Phase 4: Monitor and retrain (ongoing, $300 every 6-12 months)
First-time investment: ~$300 + 8-10 labor hours ROI: 25% effort drop with 200k words/month = ~$6,000 annually. Payback: 2 months.
Bottom Line¶
Custom machine translation is accessible, affordable, increasingly necessary for competitive agencies.
Processing 100k+ words/month with 10k+ high-quality segments? Train today. Clear ROI, modest effort, 6-12 month payback.
Processing under 50k words/month? Wait until volume justifies investment.
Treat your translation memory as a monetizable asset. That shift—from “we archive our TM” to “we train our TM”—is where competitive advantage lives.