OCR + Machine Translation + Post-Editing: The Complete Workflow

How to translate scanned documents using OCR, machine translation, and post-editing. Practical guide for translators and agencies: who does what, how long it takes, what errors to watch for at each stage.

Also in: RU EN UK
OCR + Machine Translation + Post-Editing: The Complete Workflow

27 pages of scanned medical documentation, English, deadline tomorrow. Full human translation would take two or three days. Run it through machine translation and you get a rough draft but mediocre quality. What’s the middle ground? It’s MTPE - machine translation plus post-editing. This workflow has become the economic engine for individual translators and small agencies over the past three years. Let’s break it down: how it actually works, how long it takes, and what gotchas wait at each step.

What is the OCR + MT + post-editing workflow

It’s three sequential phases, each one separate work:

  1. OCR (optical character recognition) - a scan becomes text. The computer “reads” an image, recognizes letters, outputs a text file. Accuracy depends on scan quality, font size, and font complexity.

  2. Machine translation - text from the scan goes into a model (GPT, DeepL, specialized engines), you get a draft translation. The model doesn’t verify whether OCR read correctly - it just translates what it receives.

  3. Post-editing - a translator (or editor) reviews the MT draft, compares it to the source, fixes errors, refines style. The bet is simple: if MT did okay, editing is faster than translating from zero.

These three phases together form the workflow. Let’s dig into each.

Phase 1: OCR - from image to text

A scan is essentially a large PDF or JPG image with text written on it. The computer doesn’t “see” it as text. An OCR engine analyzes pixels, recognizes characters, assembles them into words, outputs a text file. Clean scan = nearly 99% accuracy. Old, worn scan with stains = accuracy drops to 80-95%.

Factors that affect OCR accuracy:

Resolution: as noted on DocuClipper, 300+ DPI (dots per inch) is the baseline for reliable recognition. At 150 DPI accuracy drops noticeably. Below 100 DPI expect frequent errors. For small fonts or complex typefaces, use 400+ DPI.

Original condition: stains, creases, skew, poor lighting - all reduce image quality, and thus OCR accuracy.

Fonts: standard geometric fonts (Arial, Times) are easy. Decorative or ornate typefaces are harder. Handwriting is almost impossible for OCR.

Language: for Western European languages (English, German, French) OCR has been >99% for years. For non-Latin scripts (Arabic, Chinese, Devanagari) accuracy is lower.

What commonly breaks OCR:

Tables - table structure recognition often fails with cell misidentification. Math symbols and formulas - often read as garbage (“1×” instead of “$1”). Handwriting - recognition rate near zero. Very old or degraded documents.

Tactic: how to prepare a scan for quality OCR

Scan at 300+ DPI. Square up the document so it’s parallel to the camera. Make sure lighting is good (no shadows, no glare). If it’s handwriting or a very old document, decide upfront whether OCR will help or if you need human translation.

After OCR, spot-check critical areas: numbers, dates, names, coefficients. If OCR misread ‘2024’ as ‘2O24’ (zero as letter O), you’ll see it and should fix it before machine translation.

Phase 2: Machine translation

Once text from the scan is clean, it goes to machine translation. According to research on Phrase, modern models (GPT-4, DeepL, specialized) produce drafts at 70-85% quality for many language pairs.

What MT does well:

On language pairs with abundant parallel training data (EN-DE, EN-FR, EN-RU), MT produces acceptable drafts. Structure is preserved: two paragraphs stay two paragraphs. Names (from context) often translate correctly. Terminology can still drift.

What MT does poorly:

On pairs with less training data (EN-LT, EN-EL, many Asian pairs), quality suffers. Long complex sentences often translate uncertainly. Calculations - forget it. Subtle wordplay, sarcasm, puns - aren’t caught. Cultural references needing context often fall apart.

The core trap of MT:

MT doesn’t fact-check itself. If OCR misread ‘2O24’ and fed that to machine translation, MT won’t say “hey, that’s weird”. It’ll just translate it.

If the source text contains a logical error (say, “male patient with pregnancy”), MT won’t catch and fix it - it’ll translate the error faithfully.

If dates are misread or missing in input (e.g. a table of dates), the output table will have wrong dates.

Tactic: how to set context for MT

Before sending text for translation, give the model context: briefly say what the document is about, who the audience is, what tone it should have. Crowdin recommends preparing a glossary - a list of terms in the document and their translations. “Appraisal = performance evaluation”, not just “evaluation”. MT will stick to this glossary. It makes a real difference.

Remove obvious typos from the source before MT (spot ‘2O24’, fix to ‘2024’). Don’t expect MT to correct things for you.

Phase 3: Post-editing - how editors work

The draft is ready. An editor sees what MT produced, with source text alongside. Editor compares, finds errors, fixes them.

What an editor does (light PE):

Fixes obvious errors: if MT wrote gibberish, fix it. If a sentence is completely wrong or unclear, retranslate it. Reviews terminology for consistency: if “user” got translated both as “пользователь” and “юзер”, pick one and standardize.

Light PE on 1,000 words typically takes an experienced translator about 1 hour (depends on MT quality and text complexity).

What an editor does (full PE):

Fixes everything in light PE, then polishes style, corrects grammar nuances, rewrites awkward passages so it sounds natural. Full PE is almost like re-reading the whole translation. On 1,000 words: 1.5-2 hours.

Productivity:

According to industry data, someone doing MTPE produces: - Light PE: 4,000-8,000 words per day - Full PE: 3,000-5,600 words per day

Compare: human translation from scratch - 2,000 words per day. So light PE means someone does 2-4x more work.

But that’s averages. In practice it depends on: MT quality (40% error rate = slow PE), language pair (EN-FR gains bigger, EN-ZH smaller), editor experience (newcomers are 30% slower), document type (fixed structure faster, free prose slower).

Common traps when doing PE:

You didn’t check the scan before MT. OCR erred, MT propagated the error, the editor’s reviewing the error and might miss it (relying on source that’s wrong).

Editing without checking source. Editors must stay vigilant: if MT wrote something that sounds good but doesn’t match the original, that’s a translation error, not MT beauty.

Ignoring consistency. If text started with “user”, later must be “user”, not “юзер” or “users”, even if MT suggested it.

Throwing away poor MT. If MT error-rate is too high to post-edit efficiently, don’t force it - human translation from scratch may be faster.

Comparing approaches: human translation vs. OCR+MT+PE

The math becomes clear:

Approach Time per 1,000 words Quality Cost per word When to use
Human translation 4-5 hours 95-98% $0.10-0.25 Official docs, nuances matter, small volume
OCR + Light PE 1.5-2 hours 85-92% $0.02-0.05 Large volume, quality can relax a bit
OCR + Full PE 2-3 hours 90-96% $0.05-0.15 Official scanned doc, speed matters
Raw MT 0.25 hours 60-75% $0.01-0.03 Informal content, errors tolerable

Your choice depends on:

Source quality: clean, clearly scanned? Then OCR+MT+PE works. Stained, handwritten, very old? Maybe human translation straight up.

Volume: 5-10 pages = human. 50+ pages = OCR+MT+PE, advantage obvious.

Language pair: EN-DE, EN-FR = MT solid, gains visible. EN-LT, EN-JA, specialized pairs = MT weaker, gains smaller.

Deadline: tomorrow = human or full PE. A week out = light PE works. A month = light PE is good.

Tools and platforms for this workflow

For OCR: ABBYY FineReader (190+ languages), Adobe Acrobat Pro DC, Google Cloud Vision API, Tesseract (free, open-source).

For MT: DeepL, Google Translate, Claude, specialized models depending on language pair.

For post-editing workflow: memoQ, Trados, Phrase - CAT tools that provide editing interface for MT drafts, maintain glossaries, track productivity.

Or use a single platform that bundles all three phases: upload a scan, platform does OCR, machine translation, delivers result in an editing view. ChatsControl is one such platform: upload PDF or photo, it runs OCR, machine translation, shows result side-by-side (source and translation highlighted; click a word to edit). Built-in QA validator catches errors and inconsistencies.

Best practices for successful OCR+MT+PE

1. Scan right from the start

300+ DPI, proper lighting, document level and straight. Good scan = 70% of work done.

2. Clean OCR text before MT

Check critical spots (numbers, dates, names), remove typos. 15 minutes cleaning OCR can save an hour on editing.

3. Give MT context

Say what the document is about, what tone, what critical terms exist. Lokalise recommends this improves quality 10-15%.

4. Build a terminology glossary

If the document has specialized terms (“appraisal”, “stakeholder”, finance terms), write their translations BEFORE running MT. MT will use it.

5. Match editor to complexity

A first-time post-editor runs 30% slower. Give the tough job to an experienced editor if deadline’s tight.

6. Handle tables and formulas specially

Tables in OCR often explode (cells scramble). Try: extract table as plain text separately, translate that, paste back. Formulas are usually images anyway - don’t translate, just explanatory text.

7. Enforce consistency

One editor, one document if you can. If you must split across editors, give them a shared glossary and spot-check each editor’s work.

Common mistakes to avoid

1. Skipping scan preprocessing

You feed a dirty scan to OCR, get dirty text, MT translates the dirt, editor spends hours. Result: same time as human translation, but worse quality.

2. Picking the wrong MT engine

For EN-DE use DeepL or GPT-4, quality’s solid. For EN-LT choices are fewer - try a few, pick best, but don’t expect 95%. Underestimate this and you’re blaming editors when the problem is the model.

3. Editing without source comparison

Editor should see original and MT draft together. If editor only edits the translation (no source), they’ll miss translation errors (text sounds fine but diverges from original).

4. Ignoring language specifics

English plurals, Russian gender, Ukrainian verb aspects - MT often gets these wrong. Editors must know and fix.

5. Forcing poor MT into post-editing

If MT quality is 40-50%, don’t push it through PE - human translation from scratch is probably faster and better. MTPE works when MT is 70%+.

Checklist for launching an OCR+MT+PE project

  1. [ ] Identify language pair, check if MT is strong for it (EN-DE/FR ✓, EN-LT ?, EN-ZH no)
  2. [ ] Scan (or obtain) original at 300+ DPI
  3. [ ] Run OCR, spot-check critical areas (numbers, dates, names)
  4. [ ] Build terminology glossary (if doc is specialized)
  5. [ ] Run MT on cleaned text
  6. [ ] Sample-check MT quality (3-5 paragraphs) - if >70%, proceed; if <60%, decide if PE is worth it or go human
  7. [ ] Distribute to editors, give them the shared glossary
  8. [ ] Spot-check 25% of document before editors finish everything
  9. [ ] Final pass: terminology consistency, formatting, no missed errors

FAQ

How much time does MTPE save vs. human translation?

On EN-DE or EN-FR: 50-70%, so a day’s work becomes 3-4 hours (light PE). On harder pairs (EN-LT, EN-ZH) savings are 20-40%. Depends on MT quality.

What’s the minimum acceptable MT quality for post-editing?

70% and up. If sample check shows >30% error rate, translate it manually.

Should a post-editor know the source language?

Ideally yes, at least at comprehension level. If an editor doesn’t know English and is editing an English→Russian translation, that’s a problem (can’t verify accuracy).

How do I organize post-editing across a 3-5 person team?

Split the document by editor so each handles one content type (one does tables, one does prose, one does lists). Give everyone a shared terminology glossary to unify terms. Spot-check 10% of each editor’s work for consistency.

If the document has handwriting, what do I do?

OCR on handwriting performs poorly, 30-50% accuracy. Option 1: fix handwriting manually before OCR (if possible). Option 2: translate the document manually, skip OCR. Option 3: treat handwriting as image, insert without translation.

Will an embassy accept a translation made with MT + post-editing?

For certified translation, the sworn translator’s signature and final quality matter; method doesn’t (ML, human, hybrid all okay). For non-official translation, there are no requirements, so even raw MT technically passes. For semi-official (e.g. registry extract), check the specific institution’s rules.

What if MT mistranslates a table?

If the table has numbers, dates, coefficients, OCR often loses or mangles them. Try: extract table separately as text (like CSV), translate that, paste back. Or just fix the translated table manually.

Try ChatsControl

AI platform for professional translators

Try for free →