OCR for PDFs: How to Extract Text from Any PDF (Complete Guide)

Hands reviewing printed documents next to a laptop, illustrating OCR text extraction from PDFs

OCR (optical character recognition) turns a scanned PDF — which is really just a stack of pictures — into a document with real, selectable, searchable text. In this guide you'll learn how to tell whether your PDF even needs OCR, then walk through four proven ways to run it: Adobe Acrobat, free online converters, the open-source ocrmypdf/Tesseract stack on your own machine, and cloud document AI APIs. We also cover how to preserve the original layout, batch-process entire folders, squeeze out maximum accuracy across languages, and when a modern LLM-based document AI model is worth the extra cost over classic OCR.

Disclosure: some links in this guide point to commercial software. If you buy through them, we may earn a commission at no extra cost to you. It never affects our recommendations.

Table
  1. First: Is Your PDF Scanned or Native?
  2. Method 1: Adobe Acrobat Pro (Easiest for Occasional Use)
  3. Method 2: Free Online OCR Tools (Fast, but Mind Your Data)
  4. Method 3: Tesseract + ocrmypdf (Free, Local, Automatable)
  5. Method 4: Cloud OCR and Document AI APIs (Best Accuracy at Scale)
  6. How to Keep the Layout Intact
  7. Batch OCR: Processing Hundreds of PDFs
  8. Accuracy and Languages: Getting the Most Out of Any Engine
  9. Classic OCR vs LLM-Based Document AI: When to Use Which
  10. FAQ: OCR for PDFs
    1. How do I know if a PDF needs OCR?
    2. What is the best free OCR for PDFs?
    3. Does OCR change how my PDF looks?
    4. Can ChatGPT or Claude do OCR on a PDF?

First: Is Your PDF Scanned or Native?

Not every PDF needs OCR. There are two fundamentally different kinds of PDF, and the right workflow depends entirely on which one you have:

  • Native (digitally created) PDFs are exported from Word, LaTeX, a browser or a design tool. The text inside them is already text — encoded characters with fonts and positions. You can extract it directly with any PDF reader, pdftotext, or a Python library like pdfplumber. Running OCR on a native PDF is not just unnecessary; it usually degrades quality, because you replace perfect embedded text with a recognition guess.
  • Scanned (image-based) PDFs come from scanners, photocopiers, or phone cameras. Each page is a bitmap image. There is no text layer at all — to a computer, the words are just pixels. These are the PDFs that need OCR.

The 5-second test: open the PDF and try to select a word with your cursor. If you can highlight individual words, it's native (or already OCRed). If your selection draws a rectangle over the whole page like an image, it's scanned. On the command line, pdffonts file.pdf gives you the same answer: an empty font list means there's no text layer.

Watch out for the third case: hybrid PDFs, where some pages are native and some are scanned (common in contracts with scanned signature pages). Good OCR tools such as ocrmypdf detect this and only OCR the pages that need it.

Method 1: Adobe Acrobat Pro (Easiest for Occasional Use)

Adobe Acrobat Pro has had built-in OCR for decades, and it remains the most polished point-and-click option. Its "Recognize Text" feature adds an invisible text layer on top of the scanned image, so the document looks identical but becomes searchable and copyable.

Step by step:

  1. Open the scanned PDF in Acrobat Pro.
  2. Go to All tools → Scan & OCR (older versions: Tools → Enhance Scans).
  3. Choose Recognize Text → In this file.
  4. In Settings, pick the document language (critical for accuracy) and the output style. "Searchable Image" keeps the original scan visible with hidden text underneath — the safest choice. "Editable Text and Images" replaces the scan with reconstructed fonts, which looks cleaner but can distort tricky layouts.
  5. Click Recognize Text, then save. Use File → Export To if you want Word, Excel or plain text output.

Acrobat can also batch-OCR: Recognize Text → In multiple files lets you drop in a folder. The downsides are price (subscription only) and limited automation — for pipelines, the command-line tools below are far better. If Acrobat isn't the right fit for your budget or workflow, our roundup of the best OCR software compares the strongest desktop, open-source and API alternatives in one place.

Method 2: Free Online OCR Tools (Fast, but Mind Your Data)

For a one-off, non-sensitive document, browser-based converters are the fastest route: no installation, drag, drop, download. Reliable options include Smallpdf, iLovePDF, and Google Drive itself — upload a scanned PDF to Drive, right-click and choose Open with → Google Docs, and Google runs OCR on it automatically (surprisingly good for clean scans, weak on complex layouts).

Typical workflow with any online tool:

  1. Upload the PDF (most free tiers cap at 10–50 MB or a page limit).
  2. Select the document language if the tool offers it.
  3. Choose output: searchable PDF, Word, or plain text.
  4. Download the result and verify a few pages — free engines vary wildly in quality.

Two serious caveats. First, privacy: you are sending your document to someone else's server. Never upload contracts, medical records, IDs or anything confidential to a free converter; check the retention policy even for paid ones. Second, scale: free tiers throttle you quickly. If you OCR documents regularly, a local tool or an API pays for itself in the first week.

Method 3: Tesseract + ocrmypdf (Free, Local, Automatable)

For anyone comfortable with a terminal, this is the best cost/quality/privacy balance available. Tesseract is the open-source OCR engine originally developed at HP and later maintained by Google; it supports 100+ languages and, since version 4, uses an LSTM neural network recognizer. ocrmypdf is a Python wrapper that handles everything Tesseract doesn't: PDF parsing, image preprocessing, and writing a proper searchable PDF/A with the text layer perfectly aligned under the original image.

Install (macOS with Homebrew / Debian-Ubuntu):

brew install ocrmypdf  or  sudo apt install ocrmypdf tesseract-ocr-eng

The basic command:

ocrmypdf input.pdf output.pdf

That alone detects scanned pages, runs OCR in English, and produces a searchable PDF. The flags you'll actually use:

  • -l eng+deu+spa — recognize multiple languages in one document (install each language pack, e.g. tesseract-ocr-deu).
  • --deskew — straighten pages scanned at a slight angle; one of the biggest single accuracy boosts.
  • --clean — despeckle dirty scans before recognition (uses unpaper).
  • --rotate-pages — auto-fix pages scanned upside down or sideways.
  • --skip-text — skip pages that already contain text: exactly what you want for hybrid PDFs.
  • --redo-ocr — replace a bad existing OCR layer (e.g. from an old scanner) with a fresh one.
  • --output-type pdfa — produce archival PDF/A, the default and ideal for long-term storage.

A realistic production command for a messy multilingual scan:

ocrmypdf -l eng+fra --deskew --clean --rotate-pages --skip-text scan.pdf searchable.pdf

If you only need raw text rather than a searchable PDF, plain Tesseract works directly on images: tesseract page.png output.txt -l eng. For PDFs, convert pages to images first with pdftoppm -r 300 input.pdf page -png — 300 DPI is the sweet spot for recognition accuracy.

Method 4: Cloud OCR and Document AI APIs (Best Accuracy at Scale)

When accuracy matters more than cost — think invoices, handwriting, low-quality faxes — the managed APIs from the big cloud providers outperform Tesseract, especially on degraded input. The main players:

  • Google Document AI (and the simpler Cloud Vision API): excellent general OCR plus specialized parsers for invoices, receipts and forms.
  • Amazon Textract: strong on tables and key-value form extraction, returned as structured JSON.
  • Azure AI Document Intelligence: prebuilt models for common document types and trainable custom models.

The workflow is the same everywhere: upload the PDF (or point at cloud storage), call the API, get back JSON with the recognized text, word-level coordinates, confidence scores, and — this is the differentiator — structure: tables as tables, form fields as key-value pairs, checkboxes as booleans. Pricing is per page, and the specialized parsers cost several times more per page than plain OCR. For a few hundred pages a month the cost is trivial; at millions of pages it becomes a real budget line, which is when teams often move the easy documents to Tesseract and reserve the API for the hard ones.

How to Keep the Layout Intact

The biggest complaint about OCR is not misread characters — it's scrambled layout: two-column articles read across columns, tables collapsed into word soup, headers merged into body text. Three rules keep your output faithful:

  • Prefer a searchable PDF over text extraction when appearance matters. The invisible-text-layer approach (Acrobat's "Searchable Image", ocrmypdf's default) never touches the visual page — layout preservation is perfect by construction, because the original image stays on top.
  • Use layout-aware output formats for editing. If you need Word or HTML, use a tool that performs layout analysis: Acrobat's export, ABBYY FineReader (still the gold standard for complex layouts), or Document AI APIs that return coordinates so downstream code can reconstruct reading order, columns and tables.
  • For tables, go structured. Classic OCR outputs a stream of words; Textract, Document AI and Azure return actual row/column cell structure. If your PDFs are full of tables, that alone justifies an API.

One more layout tip: reading order in multi-column documents depends on correct page segmentation. In raw Tesseract you can steer it with --psm (page segmentation mode) — e.g. --psm 1 for automatic segmentation with orientation detection, --psm 6 for a single uniform block of text.

Batch OCR: Processing Hundreds of PDFs

Once the single-file workflow works, scaling up is mostly a loop. On Linux/macOS, a whole folder in parallel:

find ./scans -name '*.pdf' | parallel ocrmypdf --skip-text --deskew '{}' './done/{/}'

Notes for reliable batch runs:

  • --skip-text makes the run idempotent — re-running never double-OCRs, so you can safely resume after a crash.
  • ocrmypdf already parallelizes pages within one file; when using GNU parallel across files, cap jobs with -j 4 so you don't oversubscribe CPU.
  • Log failures: a small percentage of real-world PDFs are malformed. --force-ocr rescues some; keep a quarantine folder for the rest.
  • On Windows, ocrmypdf runs fine under WSL, or use a PowerShell loop over ocrmypdf.exe.
  • For cloud APIs, use the async/batch endpoints (Textract's StartDocumentTextDetection, Document AI batch processing) — they accept documents from object storage and are much cheaper in engineering effort than synchronous calls in a loop.

Accuracy and Languages: Getting the Most Out of Any Engine

OCR accuracy is mostly determined before recognition ever runs. In rough order of impact:

  1. Scan resolution: 300 DPI is the standard; below 200 DPI accuracy falls off a cliff, above 400 DPI you gain little and pay in file size.
  2. Deskew and denoise: a 2° tilt measurably hurts recognition. Always enable --deskew/--clean on real-world scans.
  3. Declare the right language(s): the language model constrains recognition. English-only OCR on a German document mangles every umlaut. Tesseract ships trained data for 100+ languages; combine them with + but only list languages actually present — every extra language slightly dilutes accuracy and slows the run.
  4. Scripts matter more than languages: Latin-script languages are mature everywhere; Chinese, Japanese, Korean, Arabic and Devanagari vary much more by engine. Cloud APIs generally lead on CJK and Arabic; for Tesseract, use the _vert models for vertical Japanese.
  5. Handwriting is its own problem: classic Tesseract is poor at it; Textract, Document AI and modern vision-language models handle print-style handwriting reasonably and cursive inconsistently. Always check confidence scores.

Measure, don't guess: OCR quality is scored as character error rate (CER) and word error rate (WER). A clean 300 DPI office scan should come in under 1–2% CER with any modern engine; if you're seeing worse, fix the input images before switching engines. And a workflow note: if your end goal is translating the recognized text, OCR quality caps translation quality — garbage characters in, garbage sentences out. Pair your OCR step with one of the best machine translation tools and keep the searchable PDF as your source of truth.

Classic OCR vs LLM-Based Document AI: When to Use Which

The newest option is skipping the classic OCR pipeline entirely and sending page images to a multimodal LLM (GPT-4o-class models, Claude, Gemini) or an LLM-powered document service. These models don't just transcribe — they read: they can return clean Markdown with headings and tables, answer questions about the document, normalize dates and totals, and shrug off layouts that break rule-based analysis.

Use classic OCR (Tesseract/ocrmypdf, Acrobat) when you need: exact verbatim transcription with per-word coordinates, a searchable PDF that preserves the original page, predictable cost at high volume, on-premise processing for confidential material, or deterministic output for compliance. Classic OCR never "helpfully" rewrites anything — a property you learn to love.

Use LLM-based document AI when you need: extraction of specific fields into structured JSON ("vendor, date, total, line items"), documents with wildly varying layouts, mixed handwriting and print, or downstream reasoning (summarize this 80-page contract). The trade-offs are real: hallucination risk on low-confidence regions (an LLM will confidently invent a plausible digit where classic OCR outputs a low-confidence flag), higher per-page cost, and no native concept of a searchable PDF text layer.

The pattern winning in production in 2026 is hybrid: run classic OCR to get the verbatim text layer and coordinates, then feed that text (or the image plus text) to an LLM for structuring and validation. You get auditability from OCR and intelligence from the model. For a deeper look at the tooling landscape on both sides, browse our Language & Document AI section.

FAQ: OCR for PDFs

How do I know if a PDF needs OCR?

Try selecting text with your cursor. If words highlight individually, the PDF already has a text layer and doesn't need OCR. If the selection acts like a rectangle over an image, it's a scanned PDF and needs OCR. On the command line, pdffonts file.pdf returning no fonts confirms there's no text layer.

What is the best free OCR for PDFs?

For local, private, unlimited use: ocrmypdf with the Tesseract engine — it's open source, supports 100+ languages, and produces professional searchable PDF/A files. For a zero-install one-off on non-sensitive documents, Google Drive's built-in OCR (Open with → Google Docs) is a solid free option.

Does OCR change how my PDF looks?

Not if you use the searchable-image approach (the default in ocrmypdf and Acrobat's "Searchable Image" setting): the original scan stays visible and an invisible text layer is placed underneath it, so the file looks identical but becomes searchable and copyable. Only "editable text" conversion modes redraw the page and can alter its appearance.

Can ChatGPT or Claude do OCR on a PDF?

Yes — multimodal LLMs can transcribe scanned pages, often with excellent results on hard layouts and handwriting. But they can hallucinate plausible-looking text where the scan is unclear, cost more per page, and don't produce a searchable PDF. For verbatim, auditable transcription use classic OCR; use LLMs when you need structured extraction or reasoning over the content, ideally validated against an OCR baseline.

Recommended:

Go up

This web uses cookies More info