Scanned PDF to Text: Copy Text From a Scan (Beta)

Get the text out of a scanned PDF as a plain text file, page by page, and a searchable copy of the PDF.

Opens: PDF, JPG, PNG, TIFF, WebP, BMP, HEIC, AVIF
Saves: PDF, TXT

Read the text

Scanned PDFs or pictures
Choose files… or drop them here
.pdf, .jpg, .jpeg, .jpe, .jfif, .png, .tif, .tiff, .webp, .bmp, .heic, .heif, .hif, .avif · up to 200 MB each, 200 files max
    Add more files

    PDFs are read page by page and each gives a searchable PDF. Pictures (photos, screenshots, scans; several or a multi-page TIFF) become one PDF, in the order you add them.

    Options

    Choose the language the text is written in: letters with accents and whole words are read much better.

    For documents that mix two languages, such as English terms in a Russian letter. Reading takes about a quarter longer.

    Each page's text direction is checked, and pages scanned the wrong way round are turned. The page pictures themselves are not changed.

    Pages scanned at a slight angle are straightened. Their pictures are then saved again, so the PDF can get larger.

    Typed PDFs and pages read before have text you can select already.

    More options

    Leave empty for all pages, or write pages like 1-3, 5, 8- (8 to the end). The other pages stay in the PDF as they are.

    Only for pictures: how they become PDFs.

    Your files are deleted after processing.

    Copying from a scanned PDF gives nothing, or a single picture, because there is no text in it to copy. To get the words out, the pages have to be read first: that is what OCR does.

    How to get the text out of a scanned PDF

    1. Add the scanned PDF above.
    2. Choose the language of the text.
    3. Press Read the text. The recognised text is shown on the page, and you can download it as a .txt file together with a searchable copy of the PDF.

    The text file is UTF-8, with every page's text in reading order and a page break between pages, so a word processor shows each page on its own. Paste it into Word, Google Docs or an e-mail, or open it in any text editor.

    Text file or searchable PDF?

    The text file is best when you want to edit or reuse the words: quoting a letter, filling a form again, translating a document, or feeding it to another program. Layout is not kept: columns become one after the other, and tables become lines of words.

    The searchable PDF keeps the page exactly as scanned and puts the text behind it, so you can search and copy while still seeing the original, with its stamps, signatures and pictures.

    Getting good text

    • Scans at 200 to 300 dpi are read best. Phone photos of pages work when the page fills the picture, is in focus and is photographed straight on.
    • Pick the right language; for a document that mixes two, add the second one.
    • Check numbers, names and amounts against the scan: a smudged 8 can become a 3.
    • Only pages chosen under Pages to read (in More options) are read; leave it empty for all.
    • Pages that already have text keep it, and that text is part of the text file too.

    More about what this tool reads and writes: OCR PDF.