How to Quote a Scanned Document When Word Count Finds Nothing

A scanned PDF may have no selectable text for a word counter to count. Learn how to use OCR, check the result, and quote the work transparently.

Also in: EN UK RU
How to Quote a Scanned Document When Word Count Finds Nothing

A client sends a PDF, the pages look readable, and the word-count tool reports nothing. The file isn’t necessarily empty: the PDF may contain page images rather than selectable text, leaving the counter with no text layer to inspect.

That gap matters before anyone translates a line. A quote based on an invented exact count can set the wrong fee, create a dispute about scope, or miss text hidden in a table or footnote. The practical fix is to identify what kind of PDF you have, try OCR (optical character recognition, software that turns text in an image into editable text), and check the extracted text before using its count.

A scanned-document quote can still be clear when OCR fails. State what you checked, whether the count is exact or estimated, what pricing basis you propose, and what would trigger a revised quote.

Why a word counter sees nothing in a scanned PDF

A scanned PDF can be a collection of page images rather than a document with machine-readable text. A text-based counter reads characters and words in a text layer; it can’t count the letters visible only as pixels in a page image. Google describes its Drive workflow as converting an image or PDF to text through OCR in its undated help guidance.

That distinction explains why two PDFs that look almost identical on screen can behave differently. One might let you select a sentence and copy it. Another might show the same sentence clearly but select the whole page as an image, or select nothing at all. A word counter can work on the first file and return a blank result for the second.

A file can also be mixed. Some pages may contain digital text while others are scans, or a scan may have an incomplete text layer left by an earlier conversion. A successful count of part of the document doesn’t prove that every page has been counted.

Start by testing the file, not by trusting the word-count box:

  1. Open the PDF in a viewer and try to select a visible word.
  2. Search for a word you can see on the page.
  3. Copy a short passage into a text editor and check whether it contains the intended words.
  4. Repeat the test on pages with tables, multiple columns, stamps, or faint print.
  5. If selection and search fail, or the copied text doesn’t match the page, treat OCR as the next step rather than treating the file as text-ready.

A useful everyday example is a two-page form. The first page may be digitally generated, so its text can be selected and counted. The second page may be a photographed signature page, so the same counter silently omits its printed labels. A partial result looks plausible, but the quote is based on only part of the source.

PDF appearance doesn’t reveal how much text the file contains internally. Zooming in can make a low-quality scan look crisp to a person, but zooming doesn’t add a text layer. Conversely, a file may have selectable text even when the page looks visually rough. Test the document’s actual behavior before deciding how to count it.

Apex Translations’ published, undated word-count guidance distinguishes scanned-content PDFs from electronically countable PDFs and describes using OCR to convert a scan into editable text before counting (guidance). The practical takeaway isn’t that every scan will produce a reliable number. It’s that the counting method should match the source.

How to create and check an OCR word count

OCR gives the counter something to count, but the first number it produces isn’t automatically a dependable quote. Recognition errors can split a word, merge separate words, mistake a letter for a number, or skip faint text. Count the extracted text only after checking whether it represents the page.

Google Drive’s undated help guidance describes converting multipage PDFs and JPEG, PNG, or GIF images to text by opening them with Google Docs. The same guidance recommends a file size of 2 MB or smaller for its described conversion workflow (Google Drive Help). Treat that as guidance for that particular workflow, not as a universal limit for all OCR software.

A simple Drive workflow looks like this:

  1. Upload the PDF or supported image to Google Drive.
  2. Open the file with Google Docs using the conversion option.
  3. Let the OCR process create a document containing the recognized text.
  4. Compare the extracted text with the original scan, page by page.
  5. Correct errors that affect the count, or record the count as an estimate if the text can’t be reliably repaired.
  6. Run a word-count tool on the checked text and preserve the original PDF as the reference.

OCR results depend on the source image. Google’s undated Drive guidance recommends text at least 10 pixels high, upright orientation, sharp images, even lighting, and clear contrast; it also names common fonts such as Arial and Times New Roman as preferable for the workflow (Google Drive Help). A rotated page, shadow across the text, faint photocopy, or cramped type can make a count less dependable.

The guidance also says that Drive detects the document language and links to supported languages. Check that the source language is supported before relying on the extracted count. A recognizer working against the wrong language can confuse characters, especially where alphabets or accents differ.

A page-by-page check is more useful than a quick glance at the first paragraph. Compare headings, body copy, numbers, names, and the edges of each page. Look closely at places where the layout carries meaning: multi-column text, lists, tables, footnotes, endnotes, stamps, and small-print notes.

Google warns that its conversion may preserve bold, italics, font size and type, and line breaks, while lists, tables, columns, footnotes, and endnotes are unlikely to be detected reliably. The help page specifically calls for checking layout-dependent content after conversion (Google Drive Help). A clean-looking output can still be incomplete if, for example, one column has been read before the other and the text order has become scrambled.

Chrome offers another option for certain users and files. Google’s undated Chrome help page says its PDF viewer can OCR scanned physical-document PDFs so the text becomes searchable and selectable. The page also says that this specific conversion is completed on the user’s device without sending the PDF to Google or third parties (Chrome Help). Keep that privacy statement attached to the Chrome workflow it describes; it isn’t a general claim about every OCR tool.

Check What you’re looking for What to do if it fails
Can you select a visible word? A real text selection rather than a box around the page image. Try OCR or inspect the PDF’s text layer with another available workflow.
Does search find visible text? Search results that match the words on the page. Treat the page as potentially image-only or incompletely searchable.
Does copied text match the page? Correct reading order, spelling, names, and numbers. Correct the text or avoid calling its count exact.
Are layout-dependent sections present? Lists, columns, tables, notes, and small print. Check those sections against the scan and record omissions.
Does the source language fit the OCR workflow? Recognition that matches the document’s language and script. Check the tool’s language support and review the output more closely.

Consider a contract with a clear cover page and a faint appendix. OCR may capture the cover accurately, then miss several lines in the appendix. Counting the extracted text without reviewing the scan produces a precise-looking number with a weak foundation. Checking the appendix can reveal that the number needs correction or that the quote should stay an estimate.

OCR is a preparation step, not a substitute for the original. Keep the scan open while you review extracted text. If the OCR output is hard to follow, preserve the uncertainty in the quote rather than hiding it behind a single number.

What to put in a quote when the count isn’t reliable

A translation quote should distinguish what the client supplied from what you can verify. If the source text is countable, give the source-word count and explain the pricing basis. If the PDF is a scan and OCR results are incomplete, call the figure an estimate and name the reason: poor image quality, unreadable sections, uncertain reading order, or layout that OCR doesn’t preserve.

Apex Translations’ undated published guidance says it prefers an electronically established source-text count when possible, because that gives a firm basis for costing before financial commitment. The same guidance says that when a reliable source count can’t be established, its quotations may instead use an estimate of the target-text word count (Apex guidance). That’s one agency’s stated policy, not a universal rule. It does show why the quote should make the basis visible.

A clear quote can include:

  • Document condition: scanned PDF, searchable PDF, mixed file, or image.
  • Counting method: selectable text, OCR followed by checking, or a manual estimate.
  • Count status: exact count from checked text, provisional OCR count, or estimate.
  • Pricing basis: source word, target word, page, hour, or project, as agreed.
  • Scope: which pages and visible elements the quote covers.
  • Uncertainty: text that can’t be read or checked, such as faint notes or handwriting.
  • Revision trigger: what changes the quote, such as missing pages or newly legible text.

A practical note might read: “The PDF contains scanned pages. I used OCR, then checked the extracted text against the original. The resulting count is provisional because the appendix has faint print and the table text may be incomplete. The quote covers the legible printed content shown in the supplied file; unreadable or omitted content may require a revised estimate.”

That wording doesn’t pretend the problem is solved. It gives both sides a shared record of what the number represents and where the uncertainty sits. If the client later supplies a sharper scan or an editable file, you can update the count using the improved source.

Don’t mix a page estimate and a word count as if they were interchangeable. A dense page of small text and a page with a heading, signature, and broad margins can take very different amounts of translation work. If you use pages as the basis, define what a “page” means for the quote or simply identify the supplied page range and document version.

The choice between per-word, per-page, per-hour, and project pricing depends on what can be measured and agreed. The source dossier doesn’t establish one universal tariff for scanned documents. ISO 17100:2015 describes requirements for core processes, resources, and other aspects of providing translation services to applicable specifications; the standard’s scope doesn’t prescribe a universal per-word tariff (ISO catalog). The pricing basis belongs in the agreement between the parties.

A fixed project quote can suit a bounded document when both sides agree on the pages and uncertainty. A time-based approach can make sense when deciphering or checking takes an unknown amount of work. A per-word quote may be workable after a source count is established, or when the parties explicitly accept an estimate. Whichever basis you choose, don’t label an estimate as an exact source count.

A translator working with repeat clients can also separate counting from translation in the workflow. First assess the file and report whether the count is verified. Then agree how to handle unclear pages. That step prevents the first draft from turning into a debate about what the original quote included.

The broader translation pricing guide explains how per-word, per-page, and per-document models differ. For a scan, the key question comes before the rate: can the source be counted consistently enough to apply that model?

When OCR isn’t enough, and what to do instead

OCR works best when the scan gives it clear printed text to recognize. An image can be visible and still be a poor input: faint letters, skewed pages, shadows, unusual type, or damage can create output that looks like language but doesn’t match the source. A recognizer can also capture ordinary paragraphs while losing the order of columns or text inside a table.

The word “count” can hide several different quality problems. A number can be wrong because OCR skipped words. It can be inflated because a word was split into several fragments. It can include duplicated text or omit a heading. It can be numerically close while still being unusable as a translation source because the recognized text is out of order.

Use the problem you observe to choose the next step:

  • A few clear recognition errors: correct the extracted text against the scan, then recount.
  • Repeated errors across a page: improve or replace the scan if the client can provide a better copy, then run OCR again.
  • Missing tables, columns, or notes: inspect those areas directly and don’t assume the main paragraph count includes them.
  • Text that’s visible but ambiguous: flag the uncertain passage and ask for clarification or a clearer source.
  • Handwriting or a heavily degraded image: don’t rely on automated recognition for a final count; agree on a manual estimate or a suitable review method.
  • Pages that cannot be read: name them as exclusions or quote them as uncertain work rather than silently treating them as blank.

A second OCR pass can help identify whether an error comes from the image or from a particular conversion. The output still needs comparison with the scan. Two tools agreeing doesn’t prove that either one captured a faint name or a small footnote correctly.

The scanned-document OCR workflow covers the broader role of OCR when preparing a scan for translation. For quotation, the useful boundary is simple: OCR can create countable text, but human review decides whether the count is trustworthy enough to use.

Keep three records together: the original scan, the extracted text, and the count or estimate used in the quote. Those records make later corrections easier to explain. If a client supplies a new scan, you can also distinguish changes caused by better source material from corrections to the first extraction.

A translated output shouldn’t replace the source as the authority for uncertain words. Names, identifiers, dates, and figures deserve special attention because a recognition error can affect meaning even when the total word count barely changes. A count is a costing aid; it doesn’t certify that OCR read every character correctly.

If the client needs an official translation, counting and translation quality are separate issues. A clear count doesn’t confirm that a recipient will accept the resulting translation, and OCR doesn’t make an uncertified translation certified. Ask what deliverable the client needs and keep that requirement separate from the method used to estimate the source volume.

Common quoting mistakes and a safer workflow

The most common mistake is treating the first OCR count as a fact. The number may be useful for an initial estimate, but a count from unchecked text should be labeled provisional. A single precise figure can give a false sense of certainty when pages are faint or layout-dependent content is missing.

Another mistake is quoting only from the pages that the tool happened to recognize. A report that counts page one but omits a scanned page two can look plausible, particularly when the PDF opens normally and the omitted page is visually readable. Check every page in the supplied file before calling the count complete.

A third mistake is quoting “per page” without explaining the scope. The client may assume that a blank page, a handwritten note, a stamp, or text inside a table is included. List the pages covered and identify any content you couldn’t read or count. That small clarification can prevent a disagreement after delivery.

Avoid these shortcuts:

  • Copying a count from a partial text layer: check every page, including pages that behave differently from the rest.
  • Counting raw OCR output without review: inspect names, numbers, headings, and layout-dependent sections.
  • Calling a rough estimate exact: label the method and the uncertainty.
  • Assuming visual clarity means machine readability: test selection and search rather than judging by appearance.
  • Ignoring the language setting: check the OCR workflow’s supported languages and compare the output with the source.
  • Treating a page count as a word count: define the basis and keep the two measures distinct.
  • Promising that OCR will recover everything: make the quote conditional where the scan prevents verification.

A safer workflow is short enough to use on routine jobs:

  1. Inspect the PDF. Test selection and search on several pages; note whether the document is digital, scanned, or mixed.
  2. Choose an OCR method. Use a workflow appropriate to the file and language. Google Drive and Chrome document distinct OCR options in their respective help pages.
  3. Check the input. Look for rotation, blur, faint print, uneven lighting, low contrast, and small text.
  4. Extract text. Run the conversion and save the OCR output separately from the source PDF.
  5. Review representative sections. Compare body text and any tables, columns, notes, or small-print areas against the scan.
  6. Count only checked text. Use a text editor or translation-oriented word counter on the reviewed extraction.
  7. Describe the basis. State whether the number is exact, provisional, or estimated and how the price will be calculated.
  8. Agree on exceptions. Identify unreadable content, excluded pages, and circumstances that would change the quote.

Google’s undated Drive guidance says the OCR conversion may preserve some basic formatting while missing lists, tables, columns, footnotes, and endnotes, so layout review belongs in the workflow rather than as a last-minute extra (Google Drive Help). Google’s undated Chrome guidance describes a separate local OCR workflow for scanned PDFs (Chrome Help). Use the documentation for the tool actually chosen, and don’t transfer a feature or privacy statement from one product to another.

For an internal process, a short status label can make quotes easier to compare: “verified count,” “OCR count reviewed,” or “estimate from scan.” Add a note when the document is mixed or when specific pages need manual review. The label doesn’t replace the explanation; it helps a project manager spot which quotes depend on uncertain source material.

A clear quote also leaves room for the client to improve the input. Ask whether an editable original exists, or whether they can provide a sharper scan. A Word file or text-searchable PDF may let you establish a source count directly, while a new scan may improve OCR. If neither is available, the estimate remains useful as long as both sides understand its limits.

FAQ

How do you count words in a scanned PDF for a translation quote?

Run OCR to create selectable text, review the extracted text against the scan, and count the corrected text. If OCR is unreliable, state that the count is an estimate and agree on another pricing basis or a review method before translation begins.

Can Google Drive extract text from a scanned PDF for word counting?

Yes. Google’s undated Drive help guidance describes opening a PDF or supported image with Google Docs to convert it to text (Google Drive Help). Check the extracted text before using its word count, especially where the document has columns, tables, lists, or footnotes.

How can I tell whether a PDF is scanned or text-searchable?

Try selecting a visible word and searching for a word on the page. If the viewer can’t select or find the visible text, that page may be an image without a usable text layer; Google’s Chrome help describes OCR that can make scanned PDFs searchable and selectable in its PDF viewer (Chrome Help).

What should I do when OCR cannot produce a reliable word count?

Don’t present the OCR result as an exact count. Explain what makes it uncertain, keep the original scan as the reference, and agree on an estimate, a different pricing basis, or further review.

Should scanned-document translation be quoted per word, page, or hour?

Use a basis that matches what you and the client can verify and agree. A checked source-word count can support a per-word quote; an uncertain scan may call for a clearly labeled estimate, a defined page basis, or time-based work. The quote should state which method applies and what could change it.

Try ChatsControl

AI platform for professional translators

Try for free →