How to make a scanned PDF searchable with OCR
A scanned PDF looks like a document but behaves like a photograph: search finds nothing, text cannot be selected, and any system that reads the file sees an empty page. Optical character recognition fixes that by adding real text behind the image. This guide covers checking a file, running OCR, and getting the most accurate result.
Updated September 15, 20265 min read
On this page
First, check whether the PDF is really scanned
Not every PDF you cannot search needs OCR, and running OCR on a document that already has text stacks a second, worse layer on top of a good one. Is my PDF scanned? checks each page and reports which have a real text layer and which are only images. Documents assembled from different sources often mix both.
How to OCR a PDF, step by step
Recognition runs without transmission: neither the file nor the language data is uploaded.
- If the scan is crooked, straighten it first with Deskew scans.
- Open OCR PDF and add the scanned document.
- Choose the document’s language, or several if the pages are mixed.
- Run the tool and download the result.
The wrong language turns accented characters into noise. Picking the right one matters more than any other setting.
What changes in the file
The document looks exactly as before. The original page images are kept and the recognised text is placed invisibly behind them, so words can be searched, selected and copied while nothing is redrawn.
Getting better accuracy
Recognition works line by line, so the quality of the scan drives the quality of the text.
- Clean, straight, reasonably sharp scans give the best results.
- Even two or three degrees of skew measurably lowers accuracy, which is why straightening first is worth the minute.
- Faint photocopies, handwriting and pages photographed at an angle come out worse.
- Text can be present and still be wrong, so proofread anything you will quote or rely on.