A scanned PDF opens like any other, but press Ctrl+F and nothing is found: every page is a picture. Making it searchable means reading the words on each page and adding them as text the PDF reader can find, without changing how the pages look.
How to make a scanned PDF searchable
- Add the PDF above (or several).
- Choose the language of the text, and a second one if the document mixes two.
- Press Read the text and download the searchable PDF.
Open it in any PDF reader: Ctrl+F finds words, you can select and copy sentences, and search tools on your computer or in document management systems can index it.
What changes and what doesn't
The text goes into an invisible layer placed exactly over the words of each page. The page pictures are kept as they were scanned: no recompression, no cleaning, no change in quality, so the file stays about the same size. Pages scanned sideways or upside down are turned upright by setting the page's rotation, not by redrawing it.
Pages that already have text, such as typed pages in a PDF that also has scans, are left as they are and only the scanned pages are read. If a PDF was read by another OCR program and its text is poor, choose Replace earlier OCR text: the old invisible layer is removed and the pages are read again, while typed text stays.
Pitfalls
- The language matters. A German letter read as English loses its umlauts and many words.
- Crooked scans are read, but straightening them (an option) gives better results on pages that are more than a degree or two off. It saves the page pictures again, so the file can grow.
- Password-protected PDFs can't be changed here; remove the password in a PDF program first.
- Very large pages, such as maps or posters scanned at high resolution, are kept but left without text above 100 megapixels.
More about what this tool reads and writes: OCR PDF.