How to Calculate MTPE Productivity Gains with a Pilot Project

Step-by-step guide for translation agencies: measure real MTPE performance on your content, calculate ROI, and decide whether to scale post-editing across your workflow.

Also in: RU EN UK
How to Calculate MTPE Productivity Gains with a Pilot Project

You hear managers say they switched to MTPE and saw a 150% jump in throughput. You wonder: why’s our team only at 40%? You read a benchmark article—“average 3,000-5,000 words per day”—and ask your team why they’re hitting 1,500.

The problem is simple: those web benchmarks are averages hiding huge variation. Language pair EN→FR really does get +130% faster, but EN→SV goes slower than human translation. Legal documents need extensive rework; marketing texts get slaughtered by MT. Your specific MT engine might be powerful for news but terrible for your industry jargon. Real MTPE metrics only come from a pilot test on your content, your team, and your tools.

This guide walks you through the exact process: how to design an MTPE pilot, measure results honestly, calculate ROI, and decide whether to scale. Unlike vague benchmarks, you’ll get concrete numbers for your situation—the only numbers that matter for a real decision.

Why a pilot isn’t optional—it’s essential

Let’s be direct: MTPE doesn’t pay off for everyone. Here’s a real scenario:

A translation agency handles 200,000 words annually. Someone pitched MTPE: “You’ll be 3x faster, huge savings.” The agency bought Trados, hired a post-editor, spent €2,000 on training and setup. Over a year, they pushed 150,000 words through MTPE (the rest mostly niche or small jobs). They got 20% productivity gain instead of the promised 50%. The costs didn’t break even.

Why? They tested on 500 words beforehand, then bet the farm. They didn’t know their language pair (English→Danish) is “tricky” for MT. Never checked quality on the post-edited text. Ignored that 30% of output needed rework. They burned money.

A pilot prevents that. It takes 2-4 weeks, costs a bit, but saves you from betting €10,000 on a system that won’t work.

Step 1: Choose one language pair, one domain, one PE level

Start by being specific about what you’re testing. The worst mistake is testing too many variables at once—then you can’t tell which factor caused your results.

Pick one language pair for your first pilot. EN→DE? EN→FR? Don’t test both—pick one. Languages vary wildly. EN→FR gains +130% speed according to research analyzing 90 million words across 11 pairs; EN→SV loses -7% and is actually slower than human translation. Mix them in one pilot and you can’t see which is which. You’ll end up guessing instead of knowing.

Choose the language pair where you have the most volume or the most pain. If 40% of your business is EN→DE, test that. If you’re perpetually understaffed on EN→FR and can’t find translators, test that. Your first pilot should address your most pressing commercial problem.

Pick a representative content type. What do you translate most? Technical docs? Marketing copy? Certified documents? Regulatory text? MTPE shines on technical content (consistent terminology, predictable structures, few idioms) and fails spectacularly on marketing (cultural nuance, tone, wordplay, brand voice). Full PE on marketing often ends up slower than human translation because the post-editor rewrites it anyway.

Be ruthlessly honest: if 50% of your work is marketing copy, test marketing copy, not pristine technical specs. Testing easy content to inflate your numbers is how you end up with bad ROI in production.

Full or light post-editing? - Full PE: goal is output indistinguishable from human translation. ISO 18587 compliant. Typically <2 errors per 1,000 words. For official, legal, medical, or critical content where mistakes carry legal/reputational risk. - Light PE: goal is comprehensible and usable fast. Accept grammatical imperfections, minor style issues, as long as meaning is clear. Typically 5-8 errors per 1,000 words acceptable. For internal use, technical notes, web landing pages, knowledge bases, customer support replies.

The speed difference is massive. Light PE should be 60-100% faster than full PE. Full PE might not gain much from MTPE if the content is highly specialized or creative. Research shows full PE: ~700 words/hour; light PE: ~1,000 words/hour. That’s a significant gap.

Choose based on your actual use case. Don’t fool yourself: if you’re shipping to customers, you need full PE. If it’s for internal reference, light PE wins.

Step 2: Source your test material

Minimum 5,000-10,000 words per language pair to get initial insight. If you want genuine statistical confidence, go to 50,000-100,000. More volume = clearer trends and more defensible results.

Why not 500 words? Because 500 is statistical noise. A post-editor on a tiny sample has fresh enthusiasm, hasn’t settled into rhythm, tries tricks they wouldn’t normally use, adapts rapidly. You’re measuring the novelty effect, not real performance. At 10,000 words you’re seeing how they actually work. At 50,000 words you’re seeing their real ceiling and floor—how fast on good MT output, how slow on bad output.

Use material that’s truly typical for you, not cherry-picked easy cases. Don’t select simple text where you’re already fast because you’ve translated similar content 50 times. Don’t pick pathological nightmare edge cases with 15% unknowns and zero terminology matches. Representative. Take three random jobs from last month that represent your typical turnaround. That’s your test set.

This matters because if you test on “easy” content, your MTPE numbers will look great—but production will disappoint. When you scale, you’ll hit normal volume with mixed difficulty and wonder why results tanked.

Document everything upfront so you can reproduce it later: - Start date and end date - Source text (save the full original file) - Language pair and target language variant (EN→DE is not the same as EN→AT) - Content domain and subdomain (Technical→Software Documentation vs Technical→Medical Device) - Post-editor name(s) and their experience level (junior/mid/senior) - Which MT engine and version (DeepL Jan 2026? Google Translate? Claude 4.6?) - Any domain-specific terminology lists or glossaries you’ll use - Target quality level (full PE or light PE)

This metadata is your audit trail. When someone asks in 6 months why MTPE isn’t working, you’ll trace back to the pilot specifics and remember why you made decisions.

Step 3: Run machine translation

Translate your test set with your chosen MT engine.

MT engine choice makes or breaks MTPE. Bad MT always gives bad post-editing results. If your test fails, you won’t know if MTPE is inherently flawed for your use or if your engine is just weak for your domain. That ambiguity is expensive.

Test with the engine you plan to use in production. If you haven’t decided yet, test at least 2-3 because results differ radically: - DeepL (best quality for European language pairs, EN↔DE, EN↔FR, fast, pricey) - Google Translate (largest language coverage, lower cost, quality varies by pair) - Claude/GPT (superior context handling for specialized domains, expensive per word, slower batch processing)

Don’t waste weeks on “best engine” analysis in a pilot. Pick your most likely candidate and test it. You can always re-run on a different engine later if results disappoint.

One critical detail: keep an unedited copy of the raw MT output before any post-editing. You’ll compare the before and after versions to calculate edit distance. If you only save the final edited version, you can’t measure how much work went into it.

Store them clearly: source_text.docx, mt_raw_unedited.docx, mt_post_edited_final.docx. Later you’ll compare raw vs edited to see exactly what changed.

Step 4: Post-edit and measure

Now your post-editor tackles the MT output. Requirement: make it good enough by your quality standard.

Measure three things in parallel:

Time

Have the post-editor log start and end time (per document or segment if your tool allows). Focus on actual editing time, not breaks or research.

If your tool doesn’t auto-log, time it manually. Later you’ll calculate:

words per hour = edited word count / actual editing time

Edit Distance

Edit distance = the number of character or word changes from raw MT to final edited version. It shows mechanical work volume.

Most CAT tools (memoQ, Memsource, Phrase) calculate this automatically. In Word or Google Docs, manually compare versions.

Formula: edit distance (%) = (changed words / total words) × 100

Example: 1,000 words, 350 changed (additions, deletions, substitutions) = 35% ED.

Higher ED = more work. 30% ED is normal; 70% means the MT was poor.

Quality

After post-editing, this is where most pilots fail: people skip quality review. Don’t skip it.

Sample 5-10% of output (minimum 500-1,000 words) with independent review if possible—ideally someone other than the post-editor. Use a structured checklist. Ask:

  • Count: How many errors did you find?
  • Types: Grammar, mistranslation, terminology, formatting, missing content?
  • Severity: Are errors minor (extra space) or critical (wrong number, reversed meaning)?
  • Pattern: Do errors cluster (one section is sloppy) or scatter randomly?
  • Standard: Does it meet your baseline (ISO 18587: full PE < 2 errors/1,000 words, light PE 5-8/1,000)?

Common discovery: you find 8 errors in 1,000 words. That’s way above full PE standard. Now you know full PE won’t work without extensive training or a better engine.

If quality misses your bar, you have three paths:

  1. Engine quality is weak: Try a different MT system. Re-run the test (adds 1 week).
  2. Post-editor isn’t ready: They’re learning or accustomed to different standards. Spend a day training. Run again on fresh 5,000 words.
  3. Content is hard for MTPE: This domain (legal, marketing, handwritten) doesn’t suit post-editing. Accept it and leave MTPE off this work type.

Most pilots get stuck here because people skip quality review, ship bad output to clients, and then blame MTPE instead of diagnosing the real problem.

Step 5: Calculate metrics

After the pilot finishes:

Productivity (words per hour)

Full PE: editor processed 7,000 words in 10 hours = 700 wph

Light PE: same text, faster, maybe 1,000 wph

Context: human translation averages 250 wph. If PE hits 700 wph, that’s 2.8x gain.

Average Edit Distance

If you got 50% ED overall, MT already provided 50% of the final text—your post-editor only changed 50% and left 50% untouched.

Low ED (< 30%) = strong MT, little work; high ED (> 60%) = weak MT, lots of work.

Cost per word

If post-editor rate is $0.05/word or CHF 90/hour:

  • Full PE: 700 wph at CHF 90/h = CHF 0.129/word (51% of human translation rate)
  • Light PE: 1,000 wph at CHF 90/h = CHF 0.09/word

ROI calculation

This is where theory meets cash flow. Let’s walk through a realistic example.

Scenario: You’re a mid-size agency. Human translation rate: $0.20/word. You process 500,000 words annually. Your pilot showed: - Full PE: 700 wph at CHF 90/hour = CHF 0.129/word - Edit distance: 35% average - Quality: meets full PE standards - Engine: DeepL (reliable for your language pair)

Cost comparison on 100,000 words: - Human translation: 100,000 × $0.20 = $20,000 - MTPE (full PE): 100,000 × $0.129 = $12,900 - Direct savings: $7,100

But add real costs: - MT engine license (DeepL): $2,000/year - CAT tool upgrade (Memsource): $1,200/year - Post-editor training: $1,500 (one-time, Year 1) - QA spot-checking labor: ~$800/year

Year 1 total system cost: $5,500

Real Year 1 ROI on 500,000 words: - Translation cost if human: $100,000 - MTPE cost: 500,000 × $0.129 = $64,500 - Gross savings: $35,500 - System costs: $5,500 - Net ROI: $30,000 (or 545% ROI)

Breakeven point: At ~35,000 words processed, you’ve recovered the system costs ($5,500 ÷ $0.157 per-word savings = ~35,000 words). Everything after breakeven is pure profit.

But here’s the risk: on 20,000 words/year, breakeven never happens. You pay $5,500 in system costs but only save $3,140 in translation labor. Don’t implement MTPE for small volumes.

A realistic minimum for MTPE ROI: 200,000+ words/year in a single language pair where MTPE actually works (technical, not marketing; close language pairs, not distant). Below that, the infrastructure cost eats your margin.

Step 6: Verify quality independently

This step is often skipped—don’t skip it.

Take 10% of output (minimum 500 words) and have someone else review it. Questions:

  • Does it read naturally?
  • Are terms correct?
  • Missing content or meaning errors?

If quality falls short, you’re not ready. Possible causes:

  1. Engine quality: Try a different MT system.
  2. Post-editor skill: They need training for this PE level.
  3. Content mismatch: This domain just isn’t suitable for MTPE.

Step 7: Extrapolate to full volume

If the pilot succeeded, calculate production scale.

Say your pilot showed: - 700 wph for full PE - 10% error rate (acceptable for light PE) - $0.09/word cost

You process 500,000 words annually. Staffing needed:

500,000 ÷ 700 wph ÷ 240 working days = 2.98 PE staff

Compare to human translation: 500,000 ÷ 250 wph ÷ 240 days = 8.33 translators

Savings: 5.35 fewer staff. At ~$60,000/year each = ~$320,000 saved (before MT costs, tools, training).

Step 8: Budget real costs—the hidden expenses

Most people only count the MT engine license and call it done. That’s how you end up underwater. Real costs:

Item Cost Notes
MT engine license/year $1,000-$10,000 DeepL Business ~$1,200; Google API scales by usage; Claude premium ~$3,000
CAT tool (Trados/memoQ/Memsource) $500-$5,000/year Many tools require per-user licensing. Memsource Server: $3,000/year + per-seat. memoQ: $800-$3,000 per seat
Post-editor training (initial) $2,000-$5,000 External trainer or consulting; internal: time cost of senior staff teaching
Workflow setup & consulting $1,000-$3,000 Integrating MT into your TMS, modifying QA processes, documenting new workflow
QA software/plugins $500-$2,000 Automated QA tools that flag common errors in post-edited text
Project management overhead $1,000-$2,000 Tracking MTPE projects differently, reporting separately
Year 1 total ~$6,000-$27,000 Varies by tool choices and scale

Years 2+: Licenses + support, ~$3,000-$12,000 (one-time training and setup are gone).

The key insight: Year 1 costs are always front-loaded, but they don’t scale with volume. Scaling from 100K to 500K words doesn’t double your software costs—you spread the same fixed cost over more words. At 500K words/year, ROI hits 6-12 months. At 100K words/year, you may break even after Year 2. At 50K words/year, you never break even—the infrastructure cost exceeds labor savings.

Real minimum threshold for MTPE economics: ~200,000 words/year in a language pair where it actually works.

Real-world failure modes: when MTPE loses

MTPE doesn’t always win. Before you scale, know where it fails:

Difficult language pairs

EN→SV (English to Swedish) research analyzing millions of words showed -7% throughput. MTPE was slower than human translation. Why? Swedish grammar is distant from English; word order flips unpredictably. Most MT systems underperform on distant pairs because training data is scarcer.

Other risky pairs: EN→CJK (Chinese, Japanese, Korean—word segmentation fails), EN→Hungarian, EN→Finnish. Check before you bet 6 months of planning. A quick 500-word test with each engine costs an hour and saves you from a failed rollout.

Marketing and creative content

MTPE routinely beats human translation on technical specs, user manuals, data sheets. On marketing copy? Usually the opposite. Translators rewrite tone, cultural references, wordplay—ed distance exceeds 80%. You pay post-editors to undo 80% of the MT. No speed gain; often slower than hiring a native translator.

Real example: “Our revolutionary all-in-one solution transforms your workflow” becomes “Our revolutionary-all-in-one-solution transforms your workflow process.” Grammatically fine, but lifeless. A native English speaker rewrites it, which takes time. Light PE becomes full PE in practice.

Highly structured text often gets butchered by MT. Tables flip orientation, cells merge wrong, numbers move to adjacent cells, footnote anchors break. Post-editors must manually reconstruct formatting. On a 10,000-word contract with embedded tables and exhibits, you might spend 30% of edit time fixing layout, not language.

Already-fast translators

Some teams have optimized their workflows with templates, autocorrect macros, memory leverage, specialized dictionaries. MTPE adds nothing. They already hit 400 wph. They say honestly: “We’re already fast. MTPE won’t save us time or money.” They’re right. Not all workflows benefit—know yours before betting on the tool.

Scaling from pilot to production

If your pilot succeeded and ROI is positive, scale cautiously. Rushing from pilot to production is where most MTPE deployments fail.

1. Expand the pilot (weeks 5-8 of timeline)

Don’t jump straight to full volume. Add a second language pair if your business supports it. Add a second post-editor to test team consistency. Test on 50,000 words total. Key questions: - Do results stay consistent with a new post-editor? - Does the second language pair behave like the first, or is it different? - Can your CAT tool and QA process scale to higher volume without breaking?

Results diverge from pilot sometimes. One post-editor might hit 700 wph; another hits 500 wph. One language pair (EN→DE) might yield 35% ED; another (EN→IT) might be 50%. These are real—document the variation. Your production process needs buffers for this.

2. Define quality gates precisely (weeks 7-10)

Don’t say “high quality.” Define: “ISO 18587 full PE: < 2 errors per 1,000 words. Errors defined as: grammar mistakes, mistranslations affecting meaning, terminology mismatches against client glossary, formatting breaks. Acceptable: minor style variation, minor spacing issues.”

Create a one-page QA checklist. Train all post-editors on it. Sample-check 10% of their output initially (weeks 1-4 of production), then drop to 5% ongoing. Make it routine, not an afterthought.

3. Standardize the workflow explicitly (weeks 8-12)

Write down exactly how work flows: - Client sends source → you extract text → MT processes it → post-editor handles it → QA spot-checks it → you deliver final - Who has which tools? Can post-editors work offline, or online only? - How are rework and corrections logged? How do you know who made an error? - What’s the SLA for turnaround? (MTPE might be faster—communicate that) - How do you handle edge cases (images, embedded content, complex formatting)?

This sounds tedious, but undocumented processes cause most MTPE failures. Someone skips a QA check “this time,” and it becomes the norm. Writes it down. Audit adherence monthly.

4. Monitor quality continuously (ongoing)

Over 6-12 months, cost pressure mounts. Someone says, “We don’t have time for QA checks—ship it.” Quality slides. Systematic sampling (5-10% of output, ~2,000 words/week) should stay constant. Make QA a budget line, not optional.

Metric to track: error rate per post-editor per week. If Editor A averages 1.2 errors/1,000 words and Editor B averages 3.5/1,000, you’ve spotted a training gap. Intervene early.

Tools you’ll need

Minimum: - Spreadsheet to log time, word counts, edit distance - Word counter (Word, Trados, Memsource) - Diff tool (Araxis Merge, Beyond Compare) to compare versions

Better: - Memsource, memoQ, Trados with built-in ED, time tracking, QA checks - Vendor analytics from your MT provider

Optional/costly: - BLEU/METEOR/COMET automatic quality metrics (requires technical skill)

Common pilot mistakes (and how to avoid them)

Before scaling, know the traps that kill MTPE rollouts:

Mistake 1: Testing on easy content only

You test on your best-performing account’s technical docs. Results look great: 750 wph, 25% ED, high quality. You roll out to all accounts. Suddenly you hit marketing copy, legal documents, low-resource language pairs. Productivity tanks to 350 wph. Quality is below standard. You blame MTPE and kill the program.

Fix: Test on representative content. If your mix is 40% technical, 30% marketing, 20% legal, 10% other—make your test set match that distribution. It’ll show you real production behavior.

Mistake 2: Skipping quality review

You test 10,000 words, skip the QA check (“we trust our team”), and launch. Three weeks in, a client complains about a systematic mistranslation that a QA check would have caught on day 1. You spend a week fixing, lose the client, and abandon MTPE.

Fix: Always do independent QA on 5-10% of pilot output, even if it costs a day. It’s cheap insurance.

Mistake 3: Underestimating infrastructure cost

You budget $2,000 for the MT license, forget the CAT tool upgrade ($3,000), skip the training ($2,000), don’t budget for QA labor ($1,500/year). Year 1 cost is really $8,500, not $2,000. ROI breaks down.

Fix: Budget upfront. Use the cost table above. Add 20% contingency for unknown unknowns.

Mistake 4: Picking the wrong post-editor for the pilot

You assign the busiest post-editor (they say “sure, I’ll do 10,000 words over 2 weeks”). They’re distracted, rushed, rush through QA checks. You get misleading results. You roll out, hire three more post-editors expecting 700 wph, and only get 450 wph when they’re focused and careful.

Fix: Pick someone mid-seniority who has time to do the work properly. Or run the pilot with two post-editors and average results.

Mistake 5: Not documenting assumptions

You ran a pilot on EN→DE technical docs in 2024. You hit 700 wph. Now it’s 2025, DeepL updated their engine, and your results drop to 500 wph. You’re confused because you didn’t document which engine version you tested.

Fix: Log engine version, date, any glossaries or terminology lists used, quality bar defined, post-editor experience level, everything. In 6 months, you’ll want to re-run the pilot or explain changes to stakeholders.

Pre-scaling checklist

Before you commit to expansion:

  1. Is the MT engine right? Edit distance should be <50%. If >60%, try another engine.
  2. Quality on target? Independent review showed < 2 errors/1,000 words for full PE? (Or 5-8/1,000 for light PE?)
  3. ROI positive? Calculate at your actual annual volume. Breakeven > 6 months is risky; > 12 months is unacceptable.
  4. Team ready? Do post-editors understand quality requirements? Can they hit productivity consistently?
  5. Scalable? Do one person’s results match team expectations? Or does variation suggest training gaps?

If you answer “no” to any—revisit the pilot before scaling. Don’t launch an MTPE system until every box is checked.

Conclusion: From pilot data to real money

MTPE is not a binary yes/no. It’s a tool that only pays off under specific conditions: certain language pairs, certain content types, certain team structures. There’s no universal answer. Your only honest path is piloting on your actual data.

The time cost is real (2-4 weeks). The money cost is real (€500-€2,000 in setup and testing). But compare that to the cost of betting €20,000 on a system that fails. Or worse—implementing MTPE, watching quality tank after launch, losing client confidence, and spending months undoing the decision.

Here’s what happens next if your pilot succeeds:

  1. Expand: test a second language pair, a second post-editor, 50,000 words total.
  2. Standardize: define quality gates, document the workflow, train your team.
  3. Monitor: sample-check 5-10% of output continuously. Keep quality metrics visible.
  4. Scale cautiously: add volume gradually. Don’t flip your entire operation overnight.

If your pilot shows weak results, you have options. You can try a different MT engine. You can retrain your team. You can decide this content type doesn’t suit MTPE and leave it unchanged. All are better decisions than guessing.

The core principle: measure your own context, not someone else’s benchmark. Your language pair, your content, your team, your tools. Numbers that matter are the ones you generate yourself.

Run your pilot. Measure honestly. Decide by data, not hope.

Try ChatsControl

AI platform for professional translators

Try for free →