Scanned pages you can search and copy
A scan is a picture of a page: you can read it, but you can't search it for a name, select a
paragraph or copy a table out of it. OCR (optical character recognition) reads the letters on
the page and turns them into text. This tool adds that text to your PDF as an invisible layer
behind each page, so the PDF looks exactly as it did but can be searched, highlighted and
copied, in any PDF reader. The text also comes as a plain text file (.txt) to paste
anywhere.
| Upload | You get |
|---|---|
| A scanned PDF | The same PDF with a text layer, and its text as TXT |
| Several PDFs | Each one searchable, in one ZIP with their texts |
| Photos, screenshots or scans (JPG, PNG, TIFF, WebP, BMP, HEIC, AVIF) | One searchable PDF of all of them in the order you added them (or one per picture), and the text |
The pages stay as they are
The page pictures are not changed, compressed or cleaned up: what you see is still your scan. Pages scanned sideways or upside down are turned upright (the page is turned, the picture itself stays the same). Straightening crooked pages is available too, but it saves the page pictures again, so it is off unless you choose it.
Languages
Choose the language of the text: English, German, French, Spanish, Italian, Portuguese, Dutch, Polish, Romanian, Czech, Hungarian, Turkish, Swedish, Russian, Ukrainian, Greek, Chinese (simplified or traditional), Japanese, Korean, Arabic or Hindi. For documents that mix two, add a second language. The right language matters: it tells the reader which letters and words to expect, so accents, umlauts and whole words come out right.
PDFs that already have text
Many PDFs have text already: documents saved as PDF from Word, and scans read before. Pages with text are left as they are, so a PDF that mixes typed and scanned pages only gets the scans read. You can instead replace an earlier OCR text layer (typed text stays), or read every page again, which turns pages with text into pictures first.
Under More options you can choose which pages to read, like 1-3, 8-; the other pages stay
in the PDF untouched.
What works well, and what doesn't
Printed and typed text on clean scans at 200 to 300 dpi is read almost word for word. Photos of pages taken at an angle, faded copies, very small print, decorative fonts and handwriting give more mistakes, so check names, numbers and amounts before you rely on them. Pictures with a resolution far below 150 dpi are hard to read.
Password-protected PDFs are not changed: passwords are not asked for, and a PDF whose author set a permissions password against changes is left alone too. Remove the protection in a PDF program first.
Limits
Up to 1,000 pages in one job. A typical 300 dpi A4 or Letter page takes one to two seconds; a job that would take too long is refused at once with the number of pages that fit, rather than stopping halfway. Pages larger than 100 megapixels as scanned (for example A2 at 600 dpi) are kept but left without text.
How it works
The pages are read on our server with Tesseract, the open-source OCR engine, through OCRmyPDF, which places the text exactly over the words on each page. Each job runs in its own locked-down workspace.
Privacy
Uploads are deleted as soon as the job finishes, and the results after 30 minutes (or at once with Delete now). The text is not kept, shared or used for anything else.