Copying from a scanned PDF gives nothing, or a single picture, because there is no text in it to copy. To get the words out, the pages have to be read first: that is what OCR does.
How to get the text out of a scanned PDF
- Add the scanned PDF above.
- Choose the language of the text.
- Press Read the text. The recognised text is shown on the page, and you can download it as
a
.txtfile together with a searchable copy of the PDF.
The text file is UTF-8, with every page's text in reading order and a page break between pages, so a word processor shows each page on its own. Paste it into Word, Google Docs or an e-mail, or open it in any text editor.
Text file or searchable PDF?
The text file is best when you want to edit or reuse the words: quoting a letter, filling a form again, translating a document, or feeding it to another program. Layout is not kept: columns become one after the other, and tables become lines of words.
The searchable PDF keeps the page exactly as scanned and puts the text behind it, so you can search and copy while still seeing the original, with its stamps, signatures and pictures.
Getting good text
- Scans at 200 to 300 dpi are read best. Phone photos of pages work when the page fills the picture, is in focus and is photographed straight on.
- Pick the right language; for a document that mixes two, add the second one.
- Check numbers, names and amounts against the scan: a smudged 8 can become a 3.
- Only pages chosen under Pages to read (in More options) are read; leave it empty for all.
- Pages that already have text keep it, and that text is part of the text file too.
More about what this tool reads and writes: OCR PDF.