You’re pricing an MTPE job at $0.04 per word (40% below full translation rates), and the editor tells you it’s the hardest work they’ve done all month. Meanwhile, another editor cruises through a similar-length segment in half the time. Same word count, wildly different effort—and you’re paying them identically.
This is the post-editing pricing trap that 86% of translators say is getting worse. The problem isn’t MTPE itself; it’s that word-based pricing ignores difficulty entirely. A segment with a single terminology fix takes 30 seconds. Another with broken grammar, misaligned tone, and contextual omissions takes 15 minutes. Both are “100 words,” so both get paid the same. That’s not pricing—that’s arithmetic.
What changed in 2024-2025: agencies started measuring post-editing effort at the segment level. They’re scoring each piece of machine output by how hard it actually is to fix, then pricing accordingly. Not as an arbitrary discount, but as a data-driven model that protects editor income, predicts project timelines, and gives clients honest cost estimates upfront.
Here’s how to build one.
The Three Dimensions of Post-Editing Effort¶
When a translator post-edits machine output, three distinct kinds of work happen in parallel. Lumping them together gives you a blurry picture; measuring them separately reveals where difficulty lives.
Temporal effort is the simplest: total time spent on a segment, measured in seconds or minutes. A Matecat timer or Translog II keystroke logger tracks it automatically. The problem: time alone isn’t reliable. One editor works slower out of conscientiousness; another rushes but still hits quality. Time confuses pace with effort.
Technical effort counts the actual changes made—keystrokes, insertions, deletions, substitutions, shifts. If an editor rephrases a whole sentence, that’s dozens of keystrokes. If they change one word, that’s a handful. Keystroke logging via Translog II or Matecat’s edit log captures this precisely. Studies show keystroke data varies by ~50% between different MT engines, making it a sensitive indicator of how much was actually changed. Time, by contrast, only varies 13-21% for the same text—a much weaker signal.
Cognitive effort is the mental load: how much thinking was required to fix this segment? Researchers approximate this by counting pauses (moments where the editor stops typing and thinks). If an editor pauses 10 times in a segment, that segment was cognitively taxing. If they pause once, it was straightforward. This metric is harder to capture in production (requires specialized tools like Translog II with gaze tracking), but it aligns with translator experience: “That segment made me think hard.”
Here’s the catch: these three dimensions don’t always correlate. A segment might have high technical effort (many keystrokes) but low cognitive effort (“just cleanup”). Another might have low keystrokes but high cognitive effort (“I had to rewrite the entire meaning because the source was ambiguous”). This is why single metrics fail. A fair effort-based model should account for all three.
Why Word-Based Pricing Broke¶
Until recently, most MTPE jobs were priced per word: $0.03-0.06/word depending on language pair and quality tier. It’s simple, it’s predictable for the client, and it’s transparent. It also completely ignores that a segment’s difficulty has nothing to do with its length.
Consider a real case: a pharmaceutical technical document, 10,000 words. The machine translation quality is decent—the model was trained on similar documents. Most segments are 95% usable; an editor glances, fixes a terminology glitch, and moves on. Effort per segment: 20 seconds. Cost to the client: $300-600. Fair.
Now: a creative marketing campaign, 10,000 words. The machine translation is structural gibberish—it translated idioms literally, got tone completely wrong, and invented punctuation. An editor has to practically retranslate half the document. Effort per segment: 5-10 minutes. Same $300-600 cost to the client. Catastrophic underbidding on the editor’s side.
Word count doesn’t care. An edit of “fix the comma” and an edit of “rebuild this paragraph from scratch” both count as one word. That’s why in 2024-2025, the GTS State of MTPE report showed translators shifting away from word-based rates. The model ignores the actual work being done.
Hourly rates have their own problem: client unpredictability. If you quote $25/hour and don’t know how many segments will be easy vs hard, you might end up charging the client $800 for a project you initially estimated at $400. Clients hate surprises. They want to know upfront what the project costs. Hourly rates give them neither.
Editing Distance: The Model That’s Winning¶
What’s replacing word-based pricing isn’t time-based either. It’s editing distance—a metric that captures how much of the text was actually changed during post-editing.
Editing distance has multiple names: post-editing distance (PED), edit distance, or TER-based scoring. The concept is straightforward: if an editor changes 20% of the words (through deletions, insertions, substitutions, or rewording), they receive compensation as if they’d translated an 80% TM match. If they change 50%, they get paid for a 50% match. If they rewrite 100% of the segment (basically re-translating it), they get full-translation rates.
This aligns payment with actual work. An untouched segment (0% edits)? Minimal pay—essentially proofreading rates ($0.01-0.02/word). A 30% rewrite? Mid-tier MTPE rates ($0.025-0.04). A 100% rewrite? Full translation rates ($0.06-0.12).
As one LSP put it in research cited by Taylor & Francis: “If 20% was rewritten, the linguist receives the same amount of money as for an 80% match in the translation memory.” That’s editing distance in practice—and it’s becoming standard among mid-tier and large agencies.
How do you measure it? Most modern CAT tools log it. Matecat’s editing log shows each keystroke and change. Trados Studio has edit history tracking. SmartCAT, memoQ, and others do too. Extract the edit distance per segment (usually calculated as: edited_characters / total_characters * 100), and you have a pricing basis.
The payoff: editing distance correlates strongly with actual effort. A segment with 20% edits almost always took 3-5x longer than a segment with 2% edits, regardless of whether the editor was fast or slow. It’s an objective, repeatable measure.
Implementing Effort-Based Scoring in Your Workflow¶
Moving to effort-based pricing requires three steps: measure, predict, and route.
Measure: Collect Baseline Data¶
Pick a real project—20,000-50,000 words is ideal—and track it fully for one month. Ask your editors to use Matecat, Translog II, or your CAT tool’s native logging. Capture:
- Time per segment (total time from when the segment loads to when it’s finalized)
- Keystrokes or edit operations per segment
- Character-level editing distance per segment
At the end, export the data and correlate it with the source text. Did segments from low-resource languages take longer? Did technical terminology cause more edits? Did segments with fuzzy matches require more keystrokes?
You’ll discover patterns. Maybe 60% of your segments cluster in the “easy” zone (under 30 seconds, <10% edits), 30% are “medium” (30 seconds-2 minutes, 10-40% edits), and 10% are “hard” (2+ minutes, 40%+ edits). These clusters become your pricing tiers.
Predict: Use Quality Estimation¶
Before an editor touches a segment, you can predict its difficulty using quality estimation (QE). QE tools analyze the source text and machine translation, and output a score (usually 0-1, or 0-100) indicating the likelihood that the segment is hard to edit.
Tools that do this:
- Matecat includes an MTQE (MT Quality Estimator) plugin that scores each segment inline. Run it before you assign the job.
- ModernMT offers adaptive machine translation with built-in QE scores per segment.
- TAUS EPIC (the Translation Automation User Society’s quality estimation API) scores segments on a 0-1 scale and can route them automatically to light or full post-editing.
- Language Weaver (RWS) has a quality-estimation module built into its enterprise platform.
A QE score of 0.85-1.0 means “this segment is probably good, minimal edits expected” (easy tier). A score of 0.4-0.6 means “uncertain, likely needs significant work” (medium tier). A score of 0-0.4 means “this is probably broken, expect major rework” (hard tier).
Run QE before assignment, and you price the job upfront. Instead of guessing, you know: “This 5,000-word document has 52% easy segments, 38% medium, 10% hard. Based on your historical effort data, this project will cost $450-550.”
Route: Match Difficulty to Capacity¶
Once you’ve scored segments, route them intelligently.
- Easy segments (high QE score, historically <10% edits) → junior editors or high-speed freelancers. Reduced pay rate ($0.015-0.03/word). Fast throughput.
- Medium segments (QE score 0.4-0.7, 10-40% edits) → your core post-editing team. Standard MTPE rates ($0.03-0.06/word).
- Hard segments (low QE score, 40%+ edits historically) → experienced editors or subject-matter specialists. Premium MTPE or near-full-translation rates ($0.05-0.08/word), since they might require domain knowledge or creative rewriting.
This routing does three things:
- Cost predictability. You know upfront what the project will cost the client—not “around $500,” but “$487-523 depending on final edit distance.”
- Fair compensation. Editors see that difficulty varies and pay reflects it. Hard segments pay more. Easy segments pay less. No arbitrary 40% discount across the board.
- Quality control. You’re not assigning medical terminology to a junior editor. Specialists handle specialists. Result: fewer errors, faster revision cycles, higher client satisfaction.
Building Your Pricing Model: Numbers You Need¶
To make this concrete, here’s what a real pricing card might look like based on effort data:
| Effort Tier | QE Score | Estimated Edit % | Hourly Rate Equivalent | Per-Word Rate | Per-Segment Minimum |
|---|---|---|---|---|---|
| Easy (proofreading) | 0.80-1.0 | 0-10% | $15-18/hour | $0.015-0.02 | $2-3 per segment |
| Light MTPE | 0.60-0.80 | 10-30% | $18-24/hour | $0.025-0.035 | $3-5 per segment |
| Standard MTPE | 0.40-0.60 | 30-50% | $24-30/hour | $0.035-0.05 | $5-8 per segment |
| Heavy MTPE | 0.20-0.40 | 50-80% | $30-36/hour | $0.05-0.065 | $8-12 per segment |
| Re-translation (>80% edits) | <0.20 | 80-100% | $36-45/hour | $0.065-0.12 | $15-25 per segment |
This is based on real data from mid-tier European LSPs. Your numbers will differ based on language pair, domain, and editor cost—but the structure is the same: difficulty tier determines compensation.
Note the minimum charge per segment: this protects editor income on tiny projects. A 100-word document shouldn’t pay $0.50 total; the minimum ensures baseline viability.
Real-World Case: How Editing Distance Worked¶
A technical translation agency (50-person team, 15 freelancers) switched from word-based MTPE in Q4 2024. Here’s what happened:
Before: All MTPE jobs were $0.04/word, take-it-or-leave-it. Editors complained constantly that some projects paid $4/hour (for hard segments) and others effectively $60/hour (for easy segments). Turnover was 25% annually.
During: Agency measured 10 projects (128,000 words total) over two months using Matecat logging. They discovered: - 42% of segments had <10% edits (easy) - 38% had 10-40% edits (medium) - 20% had 40%+ edits (hard)
After: They restructured pricing around editing distance: - Easy segments: $0.018/word ($18 effective hourly rate) - Medium segments: $0.038/word ($28-32 effective hourly rate) - Hard segments: $0.065/word ($36-42 effective hourly rate)
Result: Total cost to clients rose ~6% on average (because hard segments now paid fairly, not subsidizing easy ones). But editor retention jumped to 94%, cycle times dropped 12% (experienced editors no longer avoiding hard jobs), and error rates fell 8% (quality tier routing meant specialists got specialized work).
“The editing distance model felt revolutionary because it’s just… math,” one editor noted. “For the first time, I could see exactly why segment 47 paid more than segment 48. It took 8 times longer. So it made sense.”
Handling the “But My Clients Want Predictable Costs” Objection¶
Clients care about one thing: knowing the cost before they commit. Editing distance seems variable—which tier a segment lands in depends on the machine output—so won’t clients be nervous?
Not if you set bounds.
Maximum charge: “If every segment required 100% re-translation, the project would cost $X.”
Minimum charge: “Even if 80% of segments need no edits, the project will cost at least $Y.”
Expected cost: “Based on our quality estimation analysis, we expect $Z, with variance ±10%.”
This gives the client a range, and your data backs it up. In the case study above, after two months of QE data, agencies could forecast project costs with 91-95% accuracy. Clients love that. It’s more accurate than word-based estimates, which are pure guesses once you factor in difficulty variation.
You can also offer a hybrid: “We’ll charge editing distance (fair to our editors), but cap the total project cost at $600 maximum”—then you absorb the upside if the MT is better than expected. Some agencies use this as a sales tactic for new clients. Once the client sees how fair the model is, they ask to remove the cap in future projects.
When Effort-Based Pricing Fails (And What to Do Instead)¶
Effort-based pricing isn’t a silver bullet. A few scenarios where it breaks down:
Tiny projects (<500 words). You can’t calculate meaningful editing distance or QE scores from five segments. Stick to flat per-project minimums ($25-50) instead. Hourly rates work here too.
Handwritten or degraded scans. OCR quality estimation fails on bad images. You can’t predict effort because the machine output is unreliable. Here, effort estimation becomes “experienced guess.” Set fixed per-page rates or per-image fees based on quality tier (good/degraded/poor).
Highly specialized domains with no training data. Medical patent translation for a niche field? Your QE model never saw similar text, so scores are inaccurate. Offer fixed per-word rates ($0.10-0.15) or hourly, since you can’t predict effort.
Client refuses to specify domain/terminology. You can’t score effort if you don’t know the subject matter or client’s glossary. Revert to per-word or per-hour. As a condition of accepting the job, require the client to provide domain info upfront.
For all these edge cases, transparent communication is key. Tell the client: “On this project, we can’t use our predictive model because [reason], so we’re quoting [alternative pricing].”
Tools That Make Effort Scoring Easy¶
You don’t need to build effort scoring from scratch. Several platforms handle it:
Matecat ($0/month free tier + $0.10-0.25 per segment paid mode) is the most accessible. It logs every keystroke, calculates edit distance automatically, and shows post-editing effort (PEE) metrics in real time. You can export effort reports per project or per editor. For small to mid-sized agencies, Matecat alone might be enough.
Translog II (€100-500 one-time for research license, $$$$ commercial) is the gold standard for academic rigor. It captures temporal effort (second-by-second), technical effort (every keystroke), and even cognitive effort (gaze tracking with an eye-tracker). If you want bulletproof data for client negotiations, Translog II delivers it. But it’s overkill for daily operations.
ModernMT ($0-$500/month depending on usage) includes adaptive machine translation plus built-in quality estimation. You get QE scores at segment level, and you can route automatically. No keystroke logging, but QE predictions are solid for routing decisions.
TAUS EPIC ($200 per million characters, accessed via API or web interface) is enterprise-grade QE. If you’re processing 10M+ characters monthly, it pays for itself in better routing decisions. Scores are reliable; integrations are available for major CAT tools.
ChatsControl (translator.chatscontrol.com) isn’t a standalone effort-estimation tool, but if you’re doing MTPE on document translation jobs (DOCX, PDF, scanned files), it does auto-QA and effort flagging through its bilingual review interface. You can see which segments the AI flagged as problematic (easier segments to spot before editing), which helps triage difficulty.
For a small LSP just starting, Matecat + spreadsheet tracking is 80% of the value at 5% of the cost. For medium agencies, Matecat + ModernMT QE is a solid stack.
FAQ: Effort-Based Pricing in Practice¶
Q: How do I explain this to editors who’ve worked at word-based rates for 10 years?
A: Show them the data. Take a past project they worked on, extract the Matecat logs or Translog data, and calculate their historical effective hourly rate per segment. They’ll see: “Segments 1-10 paid $1-2 per segment but took 2-3 minutes each = $20-40/hour. Segments 11-20 paid the same but took 10+ minutes = $2-6/hour.” Effort-based pricing fixes that injustice.
Q: What if an editor is just slow, and editing distance doesn’t capture that?
A: It doesn’t. Editing distance captures what changed, not how fast. An editor who’s slow but accurate will show the same keystroke/edit-distance data as a fast editor—they just spent more time. If you want to factor in speed separately, add an individual productivity multiplier after the first month (some editors legitimately work slower, some are just learning). But the effort score should be unchanged.
Q: Can I use editing distance to measure quality?
A: No. Editing distance is a quantity metric (how much changed), not a quality metric (how good is the result). Low editing distance doesn’t mean high quality; it means the machine output needed few changes. Sometimes that’s because the MT was good. Sometimes it’s because the editor did minimal work (bad). Pair editing distance with spot-check QA (random samples reviewed by a senior editor) to catch quality issues.
Q: Should I implement this for in-house editors or freelancers?
A: Both. In-house editors get per-project bonuses based on effort scores (easy segments = base pay, hard segments = +20% bonus). Freelancers get per-segment or per-word rates tied to difficulty. Both models are transparent, both reward editors fairly.
Q: How often should I recalibrate my effort tiers?
A: At least quarterly. New MT engines (you switch from Google to Claude), new domains (you land a pharma contract), and new editor teams all shift effort patterns. Recalibrate every 3-6 months based on fresh data.
Takeaway: Effort-Based Pricing Is the Industry Shift¶
In 2026, word-based MTPE pricing looks as outdated as paying translators per line instead of per word. It made sense in an era when machine translation was uniformly awful, and all segments were equally hard. But modern MT is segmented: some content comes out 95% usable, some 30%, some unusable. Pretending all segments have equal effort isn’t fair to editors, and it costs clients money through hidden re-works and slow cycles.
Effort-based pricing—scored by temporal, technical, and cognitive dimensions, predicted via quality estimation, and routed to the right editor—bridges that gap. It protects editor income, gives clients honesty, and lets agencies scale without sacrificing quality.
If 86% of your freelancers are saying MTPE pricing is broken, it’s not a translator problem. It’s a pricing model problem. Fix the model, and the problem goes away.
Start small: measure one project, build one effort tier, pay one editor fairly. Then scale it. The data will convince everyone.