Poor Quality Scan Translation: What to Fix First

Learn which scan defects affect OCR and translation, when image cleanup helps, and when a fresh scan or human review is the safer choice.

Also in: EN UK RU
Poor Quality Scan Translation: What to Fix First

A tilted scan of a signed form can look perfectly readable to a person and still scramble the OCR output, turning a date into a different date or dropping a line from a table. Fixing every visible imperfection before translation wastes time; fixing the wrong one can make the source harder to read.

The practical goal isn’t to make a scan look pretty. The goal is to preserve the information a translator needs, extract text accurately enough to work with, and make uncertain passages easy to identify. Some problems respond to image cleanup. Others need a fresh scan or a human checking the source.

What needs to happen before a poor-quality scan can be translated?

A scan-to-translation workflow has three separate jobs: preserve the visual source, extract text if OCR will be used, and verify that the extracted text matches the source before translating it. Treating those jobs as one step makes errors hard to spot.

OCR, or optical character recognition, is software that turns text in an image into editable characters. OCR can save time on printed pages, but the output is not the document itself. The scanned image remains the reference for checking names, figures, dates, stamps, and any characters the software may have missed.

A useful workflow is:

  1. Keep the original. Save the source file before rotating, cropping, sharpening, converting, or recompressing it. Make a separate working copy for image changes.
  2. Check page coverage. Confirm that the full page is visible, including margins, footnotes, signatures, seals, and handwritten additions.
  3. Look for defects that change character shapes. Check focus, page tilt, shadows, glare, contrast, noise, and compression.
  4. Choose a small number of targeted corrections. Deskew visible tilt, improve uneven lighting only when needed, and avoid filters that erase faint strokes.
  5. Run OCR on the corrected copy. Keep the recognized text and image together so a reviewer can compare them.
  6. Review high-impact details against the image. Check dates, monetary amounts, document numbers, names, addresses, and negative words such as “not.”
  7. Translate only after unclear source text has been flagged. Ask for clarification or a better scan instead of filling gaps with a guess.

Microsoft lists scan quality, resolution, contrast, lighting, rotation, and text size, color, and density among the factors that affect OCR results in its current OCR characteristics and limitations guidance, checked on November 10, 2026. That list matters because a single symptom doesn’t tell you which fix will work. A faint page could be underexposed, badly printed, or simply compressed. A tilted page could also have blurred text. Diagnose the source before applying a filter.

Microsoft Learn describes the factors this way: “Document scan quality, resolution, contrast, light conditions, rotation, and text attributes such as size, color, and density can all affect the accuracy of OCR results.”

The practical reading is simple: no one image adjustment compensates for every other defect. A resolution increase can’t recreate strokes lost to motion blur, and boosting contrast can’t restore a cropped line.

One way to think about a scanned source is as a chain. First comes the image, then OCR, then human review, then translation. A defect near the beginning can travel through every later step. If OCR reads a digit incorrectly and nobody checks it, a translator may faithfully translate the wrong digit.

A scan may also be perfectly suitable for direct human reading but awkward for OCR. Conversely, OCR may produce plausible text from an image that a translator should still inspect. “The tool returned text” isn’t the same as “the source was verified.”

For a broader look at document recognition, see what AI-OCR can and can’t do with PDFs and scans. The key distinction is that OCR supports document work; it doesn’t replace checking the image that supplied the text.

Which scan defects are worth fixing first?

Start with defects that interfere with reading the actual characters or locating them in the right order. A cosmetic mark in a wide margin is usually less urgent than a shadow across a date or a tilted line that runs into the next one.

Defect Why it matters First action When cleanup isn’t enough
Page tilt OCR may misread line boundaries or group text incorrectly. Deskew a copy so text lines run horizontally. The page is curved, folded, or distorted and lines remain uneven.
Low resolution Small characters lose detail and become harder to distinguish. Return to the original scan or capture a clearer image. Fine strokes have already disappeared from the source.
Focus or motion blur Adjacent strokes merge, changing the shape of letters and digits. Check for another original or rescan with the page steady and in focus. Character boundaries are no longer visible.
Low contrast or fading Light strokes can disappear during OCR or image conversion. Preserve the grayscale source and test a restrained correction. Faint text cannot be distinguished from the background even in the original.
Shadows or uneven lighting One part of a page may be darker or lighter than another. Improve capture lighting or test a local correction on a copy. Shadows, glare, or reflections obscure characters.
Noise or compression Speckles and blocky artifacts can resemble punctuation or distort strokes. Use the least altered source available. Compression has merged or removed details.
Cropped edges Missing text, stamps, or page numbers cannot be recovered by OCR. Request a complete capture of the page. The missing area isn’t present in any source copy.

1. Deskew a visibly tilted page

Deskewing rotates a page so its text lines sit closer to horizontal. The correction can help OCR identify lines and columns, especially when the page was photographed at an angle or fed through a scanner crookedly.

Tesseract’s current quality guidance, checked on November 10, 2026, says skew reduces line-segmentation quality and recommends rotating text lines to horizontal (Tesseract documentation). That advice is specific to the engine’s OCR workflow, but the underlying issue is easy to see: slanted baselines make it harder to separate one text line from the next.

A useful check is to compare the corrected page with the original at normal reading size. Lines should look straighter, but letters should retain their shape. Avoid repeated rotating and saving because every unnecessary edit can complicate review, particularly when the file format compresses image data.

Deskewing won’t fix a page photographed with perspective distortion, where one side of the page appears larger than the other. It also won’t reconstruct a folded edge or characters hidden under a shadow. Keep those issues separate instead of expecting rotation to solve them all.

2. Restore a workable level of detail

Resolution describes how much image detail is available. OCR recommendations depend on the engine and document, so treat a specific dpi figure as guidance for a particular tool rather than a universal pass mark.

Google Document AI recommends a minimum of 200 dpi for document scans and says 300 dpi or higher generally produces the best OCR results in its current documentation checked on November 10, 2026 (supported file guidance). Tesseract’s current quality documentation, also checked on November 10, 2026, says the engine works best at 300 dpi or greater (Tesseract quality guidance). Both sources connect results to image and document quality, so meeting a resolution recommendation doesn’t guarantee accurate recognition.

A screenshot or exported PDF may not reveal the scanner’s original dpi. Don’t assume that changing a file’s metadata or enlarging its pixel dimensions creates detail. If a small character is already represented by only a few blurred pixels, stretching those pixels makes a larger blur, not a sharper original.

Small print, superscripts, footnotes, and thin table lines deserve a close visual check. Ask for a fresh capture when essential text is too small to distinguish. For existing files, test OCR on a representative page before processing a large batch; Microsoft recommends evaluating OCR against a sample representative of an organization’s own content and using confidence values to decide which results need human review (Microsoft Learn, current guidance checked on November 10, 2026).

3. Protect faded or low-contrast text

A contrast adjustment can make pale letters stand out, but aggressive processing can also erase strokes that matter. A page that’s too light may lose thin characters; a page that’s too dark may merge letters with stains, rules, or background texture.

For faded, stained, or low-inherent-contrast text, keep the grayscale image. The National Archives’ guidance for textual records says grayscale can preserve information in low-contrast, stained, or faded material and in records with handwritten annotations. Its archival recommendations specify 300-400 ppi grayscale for poor-legibility records and recommend 400 ppi in that specific archival context (NARA guidance, issued in 2002). Those figures are not a general commercial translation requirement.

The distinction between grayscale and black-and-white matters. A binary conversion forces each pixel into a black or white choice. A faint stroke that sits between those extremes may be discarded, even though a person could still make it out in grayscale.

The National Archives describes the relevant option as: “Gray scale (8-bit) scanned at 300-400 ppi.”

That line belongs to archival-transfer guidance, not a universal setting for every scanner or OCR tool. The useful takeaway is to preserve the original tonal information when text is faint, then compare any processed copy with the unaltered image.

Research on document binarization, published in 2020, describes how shadows, reflections, illumination gradients, and background distortion can cause information loss during thresholding and errors in character recognition (peer-reviewed paper). The paper reports testing a combined binarization method on 176 non-uniformly illuminated document images; that result applies to the study’s dataset, not every document or OCR workflow.

4. Reduce noise without removing punctuation

Noise is unwanted brightness or color variation in an image. Speckles can look like periods or commas; background texture can become false character strokes. Tesseract’s current guidance describes noise as a factor that can make text harder to read (Tesseract documentation, checked on November 10, 2026).

A strong cleanup filter can remove the very details a reviewer needs. Punctuation, decimal points, accents, and thin handwriting can all resemble “noise” to a filter. Use a working copy, check the result at normal reading size, and compare it with the original before accepting the change.

5. Avoid turning a scan into a different source

Repeated resizing and saving can degrade an image, especially when the format uses lossy compression. Google warns that resizing lossy formats such as JPEG to reduce file size may harm image quality and OCR accuracy in its current file guidance checked on November 10, 2026 (Google Document AI).

For translation intake, keep the original or best available master. If you need to crop or adjust contrast, create a separate working file and retain a clear link to the source. The National Archives’ 2002 rules reject lossy-compressed images for the archival transfer they cover; that is a scoped archival requirement, not a general translation rule (NARA).

What should you fix, and what should you leave alone?

The right correction depends on whether a defect changes the information or just makes the page less tidy. Use the least aggressive edit that makes the text easier to inspect, and keep the original available for comparison.

Situation Fix before OCR? Why
All text lines slope in the same direction Yes, deskew a copy. Rotation can improve line separation.
Page is faint but letters remain visible in grayscale Test a light correction, but keep grayscale. A black-and-white conversion may discard faint detail.
Page is blurry and strokes merge Usually request a clearer source. Sharpening can’t reliably restore missing boundaries.
A shadow crosses a name or amount Try a targeted correction on a copy, then compare. Broad filters can damage unaffected text.
Page edges cut off a stamp or sentence Request a complete image. Image processing can’t restore cropped content.
PDF contains clear selectable text already Check the embedded text before running OCR again. Re-OCR may introduce errors that weren’t present.
Print and handwriting overlap Preserve the image and send the passage for human review. A single OCR pass may not distinguish the layers.

A good test is whether you can answer a plain question by looking at the original: “Can I tell which character this is?” If the answer is no, a filter may create a plausible-looking mark without supplying reliable evidence. Flag the character or ask for another source.

A second test is whether the proposed fix changes all pages equally. Batch processing is efficient when a document is consistent, but pages may differ: a clean printed page can sit beside a faded page with handwritten additions. Google Document AI Enterprise OCR includes rotation correction and page-level image quality scores and defect information in its current product documentation checked on November 10, 2026 (Enterprise Document OCR). That supports a useful triage idea: inspect weak pages individually rather than treating every page as equally legible.

A document can also contain clean text plus an important exception: a handwritten correction, stamp, or signature. Don’t let the clean printed body distract from the small region that carries the legal or factual change. Mark that area for a separate human check.

Which translation workflow fits the scan?

A translation provider may work from the image directly, use OCR to extract text, or combine machine processing with human review. The best choice depends on legibility, layout, the consequence of a recognition error, and whether the output needs a formal certification.

Option 1: Ask for a new scan

A fresh scan is the cleanest solution when the source has blur, glare, missing edges, or compression that obscures characters. Request the original file if one exists, or ask for a new capture with the full page visible and text in focus.

Give clear instructions: include all page edges, keep the camera parallel to the page, avoid shadows across text, and don’t send a screenshot of a compressed preview if the original PDF is available. For a multi-page document, check that every page is included and in the right order.

A rescan isn’t always possible. Historical records, documents held by another office, and one-off photographs may have no better source. In those cases, keep the image as-is, document the unclear areas, and decide whether human transcription or a translator’s review is appropriate.

Option 2: Use OCR, then check against the image

OCR makes a scan searchable and creates editable text. It can help with printed documents, but a recognized-text layer needs proofreading against the image before translation, especially where errors could change the meaning.

Microsoft defines OCR word error rate as substitutions, deletions, and insertions divided by the number of reference words in its current guidance checked on November 10, 2026 (Microsoft Learn). The definition captures three distinct problems: a wrong character substituted for the right one, a word omitted, or a word inserted that wasn’t there.

Review should focus on more than obvious misspellings. Check names, document identifiers, dates, amounts, table entries, and words that change a statement’s meaning. A clean-looking paragraph can still contain an omitted line or a wrong digit.

Option 3: Use an online document tool for a readable scan

For a legible printed scan where layout matters, an online document-translation tool can run OCR before translation and return a formatted file. In ChatsControl, the scan workflow runs OCR first and rebuilds a clean formatted output; a bilingual side-by-side review view lets a translator compare the source and translated text (ChatsControl).

That workflow can suit a typed form or report when the source is readable and a professional can review the result. It isn’t built for handwritten or very poor-quality, degraded scans; those need a human or agency working from the original. A formatted output also doesn’t prove that OCR captured every word correctly. Check the source and translation separately.

For a scan, the current ChatsControl pricing charges at 1.3 times the character count because OCR runs first. The free tier allows 6,000 characters per month; pay-as-you-go pricing is $1.50 per 1,000 characters, or $2.70 per 1,800-character page, according to the current product pricing block. Check the live pricing before planning a job because prices and limits can change.

Online document processing is one workflow among several, not a substitute for a better source. For a certified translation, a human translator and any required certification steps still matter; the automated scan workflow doesn’t make a document legally certified.

How should an agency triage a scan before translation?

A short intake check can prevent two expensive mistakes: sending illegible material into production as if it’s clear, and asking a client to rescan a page that only needs deskewing. Record what the scan can support before quoting or assigning work.

A pre-translation scan checklist

Check Pass condition Action if it fails
Page coverage All text, margins, stamps, and relevant marks are visible. Request the missing page or a full-page capture.
Orientation Pages are upright and in reading order. Rotate or reorder a working copy.
Tilt Text lines are sufficiently horizontal to inspect. Deskew and compare with the original.
Focus Character strokes remain distinct at normal reading size. Request a sharper source if strokes merge.
Contrast Faint text and marks remain visible. Preserve grayscale and test a restrained adjustment.
OCR sample Names, dates, amounts, and headings match the image. Correct the text manually or route the page for review.
Uncertainty log Unreadable sections are recorded by page and location. Ask a focused question or request another source.

For a large batch, test a representative sample that includes the hardest pages, not only the cleanest page. Microsoft recommends using a sample dataset representative of the organization’s own content and using confidence values to identify outputs for human review (Microsoft Learn, current guidance checked on November 10, 2026). Microsoft’s example threshold of 0.80 is an example, not a universal cutoff. Choose review rules based on the documents and consequences involved.

A practical agency workflow might mark pages as ready, needs cleanup, or needs a new source or human transcription. The labels don’t need to be complicated. The point is to keep a single uncertain page from being treated as reliable because the rest of the file is clear.

Record uncertainty precisely. “Page 4, table, second row, final digit unclear” gives the project manager a useful question to send back. “The scan is bad” doesn’t tell anyone whether the issue is tilt, blur, missing content, or a hard-to-read handwritten note.

ChatsControl’s separate QA validator can help with a different part of the job: checking a translation for issues such as numbers, names, terminology, and omissions. That review doesn’t replace comparing the OCR text with the source image; it addresses translation quality after the text has been extracted and translated.

What commonly goes wrong during scan cleanup?

The most damaging cleanup errors often come from trying to make every page look equally crisp. The safest workflow preserves the original, makes targeted edits, and checks the result against the unchanged source.

Converting every page to black and white

A thresholding step can make a clean printed page look sharper, but it can also remove low-contrast strokes or annotations. The 2020 peer-reviewed paper on non-uniformly illuminated documents describes information loss during thresholding when shadows, reflections, illumination gradients, and background distortion are present (study).

Keep a grayscale copy for faded material and compare any binary version with it. If a line or mark appears only in the original, don’t assume the conversion is the truer version.

Applying the same fix to every page

A packet may include different source types: a clean typed page, a photocopy, a handwritten note, and a stamped form. One contrast setting can help one page and harm another. Review page-level problems before running batch cleanup.

Trusting a high OCR confidence score without checking

A confidence score indicates the OCR system’s estimate, not proof that the source was read correctly. Microsoft advises using confidence values to decide which outputs require human review, alongside an evaluation sample representative of the content (Microsoft Learn, current documentation checked on November 10, 2026).

A recognizer can be confidently wrong when a blurry character resembles a common alternative. Treat low confidence as a clear review signal, but don’t treat high confidence as permission to skip checks on consequential details.

Resizing a JPEG and assuming the text got better

Upscaling changes the displayed dimensions; it doesn’t recreate detail missing from the compressed source. Google warns that resizing lossy formats such as JPEG to reduce file size may degrade image quality and OCR accuracy in its current documentation checked on November 10, 2026 (Google Document AI). Keep the original and avoid repeated lossy saves.

Guessing at unreadable text

A translator shouldn’t turn a plausible guess into a source fact. Mark the unclear passage, ask for a better scan or clarification, and keep consequential uncertainties visible in the workflow. Names, dates, identification numbers, and figures deserve particular care because a single character can change the result.

The National Archives’ guidance treats OCR versions as additional to scanned images when available, rather than as a replacement for the visual record (NARA guidance, issued in 2002).

That archival practice offers a useful model for translation work: retain both the image and extracted text so reviewers can compare them. The rule itself applies to the records covered by NARA guidance; the broader lesson is to preserve a way back to the source.

A scan that remains partly unreadable after careful review needs a clear note, not a confident-sounding translation. If the unclear section affects the document’s purpose, pause that portion and ask for better evidence. If it doesn’t affect the requested translation, the translator should still make the limitation explicit rather than silently inventing missing content.

When is a new scan better than another correction?

Request a new scan when the source itself no longer contains enough visual information to distinguish the characters. Image processing can adjust what is present; it can’t reliably restore details that blur, cropping, glare, or compression removed.

Ask for a new source when:

  • Focus blur merges adjacent letters or digits.
  • A shadow or reflection hides part of a name, amount, or sentence.
  • The page edge, footnote, stamp, or signature has been cropped.
  • Small print is too indistinct to read in the original file.
  • Repeated compression has turned characters into blocky shapes.
  • OCR and a human reviewer disagree, and the original doesn’t resolve the difference.
  • Handwriting overlaps printed text and the relevant wording can’t be separated confidently.

A rescan request should name the problem and the needed action. “Please resend page 2 as the original PDF or a full-page, in-focus scan; the final digit in the amount is not readable” is more useful than “Please send a better quality file.”

When no rescan is possible, keep the source image, identify the location of each uncertain passage, and agree how to handle it before final delivery. A translator may be able to translate the legible portions while flagging gaps, but the final approach depends on the document’s purpose and the client’s requirements.

Scanned PDFs can also be turned into editable files before translation, but conversion isn’t always necessary. Compare the available routes in scanned PDF to editable Word: options before translation. A clean editable source may reduce OCR work; a poor scan converted into Word may simply carry its recognition errors into a new format.

The central question isn’t “Can software process this file?” Many tools can attempt processing. The useful question is “Can a person verify the text that matters against the source?” If not, get a better source or make the uncertainty explicit.

FAQ

What should I fix in a blurry scan before translating it?

Start with the original file and check whether the blur comes from low resolution, motion, focus, or compression. Deskew clear page tilt and test image cleanup on a copy, but request a sharper source if character strokes are merged or missing.

What scan resolution is best for OCR and translation?

Google Document AI recommends a minimum of 200 dpi, with 300 dpi and higher generally producing the best OCR results in its current documentation checked on November 10, 2026 (Google guidance). Tesseract also says its engine works best at 300 dpi or greater in documentation checked on November 10, 2026 (Tesseract guidance). Those are tool-specific recommendations, not a guarantee for every document.

Should I deskew a scanned document before OCR?

Yes, when text lines are visibly tilted. Tesseract’s current quality guidance says skew reduces line-segmentation quality and recommends rotating text lines to horizontal (Tesseract documentation, checked on November 10, 2026).

Can low-contrast or faded documents be translated reliably?

Sometimes, if text strokes remain visible in the source. The National Archives says grayscale is appropriate for low-contrast, stained, or faded textual records in its archival guidance issued in 2002 (NARA); uncertain passages still need comparison with the image and human review.

When should a translator request a new scan instead of correcting the image?

Request a new scan when focus blur, missing edges, glare, compression, or faint strokes have removed information that image processing can’t restore. A reviewer should flag genuinely unreadable text rather than infer names, dates, amounts, or other consequential details.

Does OCR confidence prove that a translation is accurate?

No. OCR confidence helps identify text that may need review, but it doesn’t prove that recognized words match the image or that a translation is correct. Compare extracted text with the scan, then review the translation separately.

Should I use grayscale or black and white for a faded document?

Keep the grayscale original and test any black-and-white conversion on a copy. NARA’s 2002 archival guidance says grayscale can preserve relevant information in faded, stained, or low-contrast textual records (NARA); that guidance is scoped to archival transfers, but the preservation principle is useful when faint details matter.

Can an online translation tool handle a handwritten or badly degraded scan?

Don’t assume it can. ChatsControl runs OCR before translating scanned documents, but the product isn’t built for handwritten or very poor-quality, degraded scans; those need a human or agency working from the original.

Try ChatsControl

AI platform for professional translators

Try for free →