A contract arrives as a PDF, but you can’t select a sentence, search for a name, or copy a figure into a spreadsheet. The page looks clear on screen, yet a document translator may see only an image. That mismatch is why a flattened PDF can produce a file that looks translated while leaving the words inside the scan untouched.
The reliable route is to check what kind of PDF you have, recognize the text if needed, review the recognition, and only then translate. The right order matters: OCR, short for optical character recognition, reads text from an image. OCR does not translate the text, check every character, or certify the result.
What makes a PDF flattened and why it changes the workflow¶
A flattened PDF is a document whose page content has been reduced to an image or image-like layer, rather than stored as selectable text. A scanned paper contract is a common example: the PDF shows the page, but the computer has no words to pass into a text-based translation workflow.
Try selecting a sentence with your cursor. If you can’t select individual words, copy a paragraph, or find a name with the PDF search function, the page may be image-only. A PDF can also contain both selectable text and scanned pages, so check more than the first page before deciding how to process it.
A selectable-text check is a useful first clue, not a full quality assessment. Text can be selectable even when the extraction order is scrambled, and a scanned page can have a hidden text layer that contains recognition errors. For a business document, inspect the words you extract rather than treating selectable text as proof that it’s ready to translate.
The format matters because a translator needs readable text, not just a picture of words. OCR creates machine-readable text from the image. A translation system can then work with that text, while a person can compare the recognition against the visible page.
The image itself doesn’t tell software where one paragraph ends or whether a group of numbers belongs in a particular table cell. OCR software has to infer that structure from the page. A skewed scan, faint print, overlapping stamp, or small label can make that inference harder.
A flattened PDF isn’t necessarily a bad scan. It can be sharp and easy for a person to read while still being unavailable to text-based tools. Conversely, an old or degraded scan may be image-only and difficult for both OCR and a human reviewer. The file type tells you what kind of processing it needs; the page quality tells you how closely to check the result.
A practical decision tree starts with the source:
- If an editable DOCX or PPTX exists, use that file as the translation source rather than converting a final PDF back into text.
- If the PDF contains selectable text, check whether copying it produces readable text in the right order.
- If the PDF contains image-only pages, run OCR before relying on a workflow that translates document text.
- If some pages are selectable and others are scans, inspect each type separately.
- If the scan contains handwriting or badly degraded print, plan for human review against the original image.
For example, a company may receive a signed agreement where typed clauses are selectable but the signature page is a scan. A translation service that handles only the selectable layer may process the clauses while missing text printed on the image-only page. Google’s Cloud Translation documentation says scanned content in a mixed PDF containing scanned and native content isn’t translated. That makes a page-by-page check more useful than relying on the PDF’s file name or extension.
The main trap is assuming that “PDF accepted” means “all words in the PDF will be translated.” File upload support and image-text recognition are separate capabilities. Check the product’s instructions for scanned pages, then inspect the result rather than judging it by whether a file downloaded.
What to do before you choose a translation method¶
Start by collecting the original materials and setting the purpose of the translation. An editable source file, if available, is usually more useful than a flattened export because it preserves the actual text and structure. Google Cloud documentation recommends translating DOCX or PPTX before converting to PDF when an editable source exists, because those formats preserve layout and style better than PDF.
Next, identify the document’s important elements. Mark names, dates, reference numbers, amounts, table headings, footnotes, handwritten annotations, seals, and text that crosses a page boundary. Those elements are easy to misread or detach from their context during OCR or translation.
Decide whether the output needs to be editable, visually close to the original, or both. A Word file is easier to correct and reuse. A page image with translated text placed over it can be useful when the reader needs to compare the translation with the original layout. Neither output should be assumed to reproduce an official document’s legal status.
Finally, check what the recipient expects. A business team may need a working translation for review, while a public authority may require a particular translator or certification process. OCR and machine translation don’t settle those requirements. If certification matters, confirm the receiving organization’s rules before choosing a workflow.
Google Translate’s help page says its document upload workflow accepts PDF, DOCX, PPTX, and XLSX files up to 10 MB, with PDF uploads limited to 300 pages, according to the help page at research time in 2026 (Google Translate Help). The same help page says document translation isn’t available on smaller screens or mobile, so its described upload workflow uses a desktop browser. Limits and interface details can change; check the linked page before preparing a large file.
Google Cloud documents a different route for document translation. Its online PDF workflow allows files up to 20 MB; native PDFs can be up to 300 pages when the native-PDF setting is enabled, while scanned PDFs are limited to 20 pages, according to the documentation at research time in 2026 (Google Cloud document translation documentation). Cloud Translation also requires a Cloud project, API access, and authentication. That setup may suit a team building a repeatable technical workflow, but it isn’t the same as opening a consumer translation page and uploading a file.
A common mistake is to start with a tool and only later discover that the file is too large, the scan isn’t supported, or the output isn’t editable. Record the file type, approximate size, page count, language pair, and required output before comparing methods. Those details make the choice concrete without assuming every PDF needs the same process.
Your options: bureau, freelancer, or an online workflow¶
A bureau can take responsibility for a translation project that needs human coordination, a formal review process, or a certified service. Ask how it handles scanned pages, whether OCR text is checked against the images, how it treats stamps and handwritten notes, and what format it returns. A quote based only on visible page count may not tell you whether image recognition or layout reconstruction is included.
A freelancer can be a good fit when you have a clear source file, a defined subject area, and a way to agree on review and delivery requirements. Share a representative sample before sending a large file. Ask whether the freelancer can read the scan, whether you should supply a DOCX version, and how they will handle table structure or text that OCR can’t read confidently.
Online document tools can reduce manual copying when you need an editable draft or need to process a document in a structured file workflow. Their capabilities differ: some accept PDFs but don’t translate scanned image text; others can recognize scans but still need a person to review the result. An online service also means uploading the document to a cloud system, so consider your organization’s data rules before using one.
A typical situation helps explain the trade-off. A team receives a clear printed form as a scan and needs a translated Word draft for internal review. OCR can turn the visible form text into editable text, but a reviewer still needs to compare names, amounts, and form fields with the original. If the same file contains faint handwriting, the automated route becomes less dependable for that section.
In ChatsControl, you can upload a scanned PDF and receive a formatted Word translation plus a copy with translated text laid over the original page image. The tool reads the scan first and rebuilds its structure for translation; its page quality grade can flag a page that needs closer checking. Handwriting and very degraded scans aren’t a good fit for automated reading, so use a human review against the original when those parts matter. ChatsControl is a document-translation option, not a full CAT tool or translation management system.
| Route | When it fits | What to check | Main limitation |
|---|---|---|---|
| Translation bureau | You need human coordination, review, or a certified service | Ask whether OCR, image checking, layout work, and certification are included | The word “PDF” in a quote doesn’t explain how scan pages are handled |
| Freelance translator | You can define the scope and agree directly on review and output | Share a sample; confirm how unclear text, tables, and images will be treated | You may need to arrange OCR or layout work separately |
| Online document workflow | You need a document-based draft and can review the output | Check scanned-PDF support, file limits, output type, privacy rules, and OCR review tools | An upload limit or accepted PDF format doesn’t guarantee image text will be translated |
| OCR followed by a separate translator | You want to control recognition and translation as distinct stages | Correct the recognized text before sending it for translation | More handoffs can mean more checking and file management |
The best comparison question isn’t “Which tool accepts PDFs?” Ask, “What does the tool do with the words that exist only as pixels?” Then ask whether the output preserves useful structure, whether you can correct OCR mistakes, and who is responsible for checking the translation.
Google Translate’s help page explains why that distinction matters:
“You can find text in images and scanned .pdf pages in the output document but they aren’t translated.”
The practical meaning is simple: a downloaded document may still contain untranslated text from scanned pages. Don’t treat the presence of the words in the output as evidence that the translation step processed them.
Google’s Cloud Translation documentation also says the scanned-PDF route can lose formatting:
“Translating scanned PDF files results in some formatting loss.”
A translated text layer and a faithful layout are separate goals. Check whether the output can preserve the elements your reader needs, especially form fields, tables, and labels.
A step-by-step workflow for a flattened PDF¶
A reliable workflow separates file assessment, OCR, translation, and review. Keeping those stages distinct makes it easier to spot whether an error came from the scan, recognition, translation, or layout reconstruction.
-
Find the best available source. Ask the document owner for the original Word, PowerPoint, or other editable file. If that file exists, use it rather than treating the flattened PDF as the only source. Google Cloud documentation recommends translating DOCX or PPTX before converting to PDF when the editable source is available (Google Cloud documentation).
-
Check the PDF page by page. Try selecting and copying text from several pages, including pages with tables, signatures, and different layouts. Paste a sample into a plain-text editor and check whether the words remain in reading order. A searchable PDF can still have a broken extraction order, and a mixed PDF can contain both image pages and text pages.
-
Prepare the scan for recognition. Use the clearest copy available. Adobe’s Acrobat tutorial describes a scan-to-searchable-PDF workflow that can enhance and straighten an image before selecting text recognition (Adobe Acrobat tutorial). Straightening and image cleanup can help the recognition stage, but they don’t replace a text check.
-
Run OCR and save a copy. OCR should produce searchable or editable text from image content. Keep the original PDF unchanged so you can compare every uncertain line with the source. If your OCR tool offers a confidence or quality indication, use it to prioritize review, not as proof that the recognized text is correct.
-
Review the recognized text against the page image. Check every high-risk field before translation: personal and company names, dates, amounts, account or reference numbers, table cells, and words covered by stamps. Compare the reading order too. OCR may recognize the right words but attach a figure to the wrong row or place a footnote in the wrong paragraph.
-
Translate the checked source. Choose a workflow that accepts the type of file you now have and returns an output you can inspect. If you use a document translation tool, confirm whether it translates scanned content or merely displays it. If you use a person, send the reviewed OCR text and the original page images so the translator can see context.
-
Review translation and layout separately. Compare meaning against the source text, then check the translated document’s tables, columns, labels, headers, and page breaks. A grammatically sound paragraph can still be attached to the wrong row, while a visually neat file can still contain a translation error.
-
Prepare the delivery copy. Keep the original scan, OCR file, translated file, and any review notes clearly separated. Label working drafts as drafts if they haven’t been approved for external use. If the recipient requires certification, confirm that step separately; OCR and translation don’t certify a document.
Adobe’s tutorial describes OCR as a recognition step:
“The file is now a searchable, editable PDF file.”
That describes a change in what the PDF’s text can do. It doesn’t say that the recognized wording is translated or that every letter has been checked. Treat the editable file as a working source that still needs review.
A common workflow failure happens when someone reviews only the final translated PDF. If a figure was misread during OCR, a translator may faithfully translate the wrong figure. Checking the OCR text against the image before translation makes that upstream error easier to catch.
How to keep tables, names, and layout under control¶
OCR and translation both affect document structure. A scanned page doesn’t contain a table as a set of editable cells; it contains visible lines, text, and spacing that software must interpret. After text extraction, a translation system must also decide where translated phrases belong. If the process turns the page into a plain stream of words, the relationship between a heading and a value can disappear.
Google Cloud warns that complex PDF layouts can lose formatting, including data tables, multiple columns, and graphs with labels or legends (Google Cloud documentation). The risk isn’t only cosmetic. A number separated from its row label may be read as a different value, and a graph legend separated from a chart may no longer identify the right series.
Build a review pass around the way a reader uses the document:
- Tables: Match each translated cell to its original row and column. Check totals, units, and merged cells separately.
- Forms: Confirm that each field label remains beside the correct blank or answer. Pay attention to text split across several boxes.
- Multiple columns: Compare the reading order with the original page. A text extractor may move to the next column too early.
- Graphs: Check translated axis labels and legends against the visual categories in the chart.
- Headers and footnotes: Confirm that page labels and notes remain attached to the correct section.
- Names and numbers: Compare letter by letter or digit by digit with the source image rather than relying on context.
Long names, similar characters, and punctuation are easy to overlook when a reviewer reads quickly. A decimal mark can change the interpretation of a figure; a single letter can change a person’s name or an account reference. OCR may also confuse a printed mark with a stamp or signature, especially where the image is crowded.
ChatsControl addresses a different part of the layout problem after recognition. Its scanned-document workflow rebuilds a formatted Word file from the text and structure it reads, and it can also place translated text over a copy of the original page image. The overlaid copy retains the original pixels for stamps and seals; the translated wording is re-typeset, so it isn’t a pixel-for-pixel original. Use the Word file when you need editable text, and inspect the overlaid copy when comparing placement matters.
No output format solves every use case. A Word document is easier to edit, but it can reflow compared with the scan. An image-based copy can keep the original page visible, but text placed over an image may need careful visual checking. Choose based on what the reviewer or recipient needs to do with the translation.
For a recurring business document, create a short review checklist using the fields that matter to your organization. A form with supplier details may need a different check from a financial table or a technical manual. The checklist should tell the reviewer what to compare, not merely ask whether the file “looks right.”
What commonly goes wrong and how to respond¶
The first pitfall is relying on the file extension. A PDF can contain native text, image-only pages, or a mixture. Google Cloud’s documentation states that scanned content in mixed PDFs isn’t translated. A file that contains a few selectable paragraphs can still leave other scanned sections untouched.
The second pitfall is treating OCR as a guarantee. OCR creates a text layer; it doesn’t promise that every name, figure, or mark was recognized correctly. Compare the extracted text against the page image, especially where the source is faint, skewed, compressed, or partly obscured. If the reviewer can’t read the original clearly, don’t silently guess at the missing word.
The third pitfall is reviewing only the translated wording. A translation can be linguistically clear while carrying forward a source recognition error. Check the OCR text first, then check the translation. That sequence helps isolate whether a discrepancy started in the scan or during translation.
The fourth pitfall is assuming that layout has been preserved because the page count looks similar. Tables can shift, columns can merge, and labels can become detached from graphs. Google Cloud explicitly warns that complex PDF layouts can lose formatting. Review content relationships, not just the appearance of the first page.
The fifth pitfall is using a consumer document upload flow on a phone and assuming the same controls are available. Google Translate Help says document translation isn’t available on smaller screens or mobile according to its help page at research time in 2026 (Google Translate Help). Use the desktop-browser workflow described in that help page, and confirm its current instructions before relying on it.
Google Translate Community includes user reports about trouble uploading scanned PDFs, including a discussion titled “Why do I keep getting a ‘Can’t translate scanned PDFs’ message” (Google Translate Community). Community posts are anecdotal, not official product specifications. The useful lesson is to distinguish an upload error from a translation result: a file that fails to upload needs a different handling route, while a file that uploads may still contain untranslated image text.
The sixth pitfall is sending an unreviewed machine translation to a recipient who needs a formal document. A useful internal draft and a certified translation are different deliverables. Ask the receiving organization what it requires before translating, and arrange any human certification as a separate step where needed.
Choosing a workflow for your document¶
Use the following checklist before you hand off a flattened PDF:
| Check | What to record | Why it affects the choice |
|---|---|---|
| Source file | Whether an editable original exists | An editable source avoids rebuilding text from an image |
| PDF content | Selectable text, scanned pages, or a mixture | A mixed file can contain text that a scan-only process misses |
| Scan condition | Clarity, skew, faint print, handwriting, stamps | Poor or handwritten material needs closer human review |
| Structure | Tables, columns, forms, charts, footnotes | Complex layout needs a deliberate formatting check |
| Translation purpose | Internal draft, business use, or formal submission | The recipient’s requirements may call for review or certification |
| Output format | Editable Word, PDF, or visual comparison copy | Different outputs support different review tasks |
| Data handling | Whether cloud upload is permitted | Online processing may not suit every organization’s rules |
Choose an editable-source route when the original document exists. Choose OCR followed by review when the only source is an image-based PDF and the text is clear enough to recognize. Choose a human-assisted route when the scan contains handwriting, poor-quality pages, or content that must be handled under formal submission requirements.
Google Cloud’s documentation says native PDFs generally offer better format handling and recommends using editable DOCX or PPTX before PDF when those sources exist (Google Cloud documentation). That guidance supports a straightforward rule: don’t make the PDF the source of truth if a cleaner editable file is available.
For an online document workflow, compare the actual sequence rather than the marketing label. Ask whether the tool reads image text, how it represents tables, what files it returns, and how it surfaces uncertain recognition. A tool that translates editable files and one that recognizes scans solve different stages of the problem.
FAQ¶
How do I translate a flattened PDF that has no selectable text?¶
Run OCR first to convert the page image into machine-readable text. Review the recognized text against the original PDF, then translate the checked text or document. OCR is recognition, not translation.
Should I run OCR before translating a scanned PDF?¶
Yes, if the PDF contains image-only text. OCR makes that text searchable or editable, but it doesn’t validate the recognition. Check names, figures, dates, and table cells before sending the text for translation.
Can Google Translate translate text inside scanned PDF pages?¶
Google Translate Help says text from images and scanned PDF pages may appear in the output document but isn’t translated (Google Translate Help). Run OCR before using a text-based translation workflow, and check the resulting text.
How can I keep tables and page layout when translating a flattened PDF?¶
Start with the editable source if you have one. Otherwise, choose an OCR workflow that recognizes page structure, then review table cells, columns, graph labels, headings, and page breaks in the translated output. Google Cloud warns that complex PDF layouts can lose formatting (Google Cloud documentation).
How do I check OCR errors before translating a business document?¶
Compare the extracted text with the page image, focusing on names, dates, figures, table cells, and text obscured by stamps. Correct recognition errors before translation so a translator doesn’t carry the wrong source text into the result.
Can I translate a flattened PDF directly in Google Cloud?¶
Google Cloud documentation supports scanned PDFs, but it warns that scanned-PDF translation can lose formatting and that scanned content in mixed scanned-and-native PDFs isn’t translated (Google Cloud documentation). Check the current file limits and prepare to review the output.
Does OCR make a PDF translation accurate or certified?¶
No. OCR reads text from an image; it doesn’t translate, validate, or certify the document. Review recognition and translation separately, and confirm the receiving organization’s requirements before using the file for a formal submission.
Should I use an online tool for a confidential business PDF?¶
Check your organization’s data-handling rules before uploading a document to a cloud service. If cloud processing isn’t permitted, choose an approved offline workflow or arrange processing through an authorized provider. Don’t assume every online tool offers the same privacy controls.