OCR PDF (Beta)

Make scanned PDFs searchable and copyable, and get the text out of photos, screenshots and scans (JPG, PNG, TIFF, HEIC and more) in 22 languages.

Opens: PDF, JPG, PNG, TIFF, WebP, BMP, HEIC, AVIF
Saves: PDF, TXT

Read the text

Scanned PDFs or pictures
Choose files… or drop them here
.pdf, .jpg, .jpeg, .jpe, .jfif, .png, .tif, .tiff, .webp, .bmp, .heic, .heif, .hif, .avif · up to 200 MB each, 200 files max
    Add more files

    PDFs are read page by page and each gives a searchable PDF. Pictures (photos, screenshots, scans; several or a multi-page TIFF) become one PDF, in the order you add them.

    Options

    Choose the language the text is written in: letters with accents and whole words are read much better.

    For documents that mix two languages, such as English terms in a Russian letter. Reading takes about a quarter longer.

    Each page's text direction is checked, and pages scanned the wrong way round are turned. The page pictures themselves are not changed.

    Pages scanned at a slight angle are straightened. Their pictures are then saved again, so the PDF can get larger.

    Typed PDFs and pages read before have text you can select already.

    More options

    Leave empty for all pages, or write pages like 1-3, 5, 8- (8 to the end). The other pages stay in the PDF as they are.

    Only for pictures: how they become PDFs.

    Your files are deleted after processing.

    Scanned pages you can search and copy

    A scan is a picture of a page: you can read it, but you can't search it for a name, select a paragraph or copy a table out of it. OCR (optical character recognition) reads the letters on the page and turns them into text. This tool adds that text to your PDF as an invisible layer behind each page, so the PDF looks exactly as it did but can be searched, highlighted and copied, in any PDF reader. The text also comes as a plain text file (.txt) to paste anywhere.

    UploadYou get
    A scanned PDFThe same PDF with a text layer, and its text as TXT
    Several PDFsEach one searchable, in one ZIP with their texts
    Photos, screenshots or scans (JPG, PNG, TIFF, WebP, BMP, HEIC, AVIF)One searchable PDF of all of them in the order you added them (or one per picture), and the text

    The pages stay as they are

    The page pictures are not changed, compressed or cleaned up: what you see is still your scan. Pages scanned sideways or upside down are turned upright (the page is turned, the picture itself stays the same). Straightening crooked pages is available too, but it saves the page pictures again, so it is off unless you choose it.

    Languages

    Choose the language of the text: English, German, French, Spanish, Italian, Portuguese, Dutch, Polish, Romanian, Czech, Hungarian, Turkish, Swedish, Russian, Ukrainian, Greek, Chinese (simplified or traditional), Japanese, Korean, Arabic or Hindi. For documents that mix two, add a second language. The right language matters: it tells the reader which letters and words to expect, so accents, umlauts and whole words come out right.

    PDFs that already have text

    Many PDFs have text already: documents saved as PDF from Word, and scans read before. Pages with text are left as they are, so a PDF that mixes typed and scanned pages only gets the scans read. You can instead replace an earlier OCR text layer (typed text stays), or read every page again, which turns pages with text into pictures first.

    Under More options you can choose which pages to read, like 1-3, 8-; the other pages stay in the PDF untouched.

    What works well, and what doesn't

    Printed and typed text on clean scans at 200 to 300 dpi is read almost word for word. Photos of pages taken at an angle, faded copies, very small print, decorative fonts and handwriting give more mistakes, so check names, numbers and amounts before you rely on them. Pictures with a resolution far below 150 dpi are hard to read.

    Password-protected PDFs are not changed: passwords are not asked for, and a PDF whose author set a permissions password against changes is left alone too. Remove the protection in a PDF program first.

    Limits

    Up to 1,000 pages in one job. A typical 300 dpi A4 or Letter page takes one to two seconds; a job that would take too long is refused at once with the number of pages that fit, rather than stopping halfway. Pages larger than 100 megapixels as scanned (for example A2 at 600 dpi) are kept but left without text.

    How it works

    The pages are read on our server with Tesseract, the open-source OCR engine, through OCRmyPDF, which places the text exactly over the words on each page. Each job runs in its own locked-down workspace.

    Privacy

    Uploads are deleted as soon as the job finishes, and the results after 30 minutes (or at once with Delete now). The text is not kept, shared or used for anything else.

    Frequently asked questions

    Does OCR change how my PDF looks?

    No. The recognised text is added as an invisible layer behind each page, so the scan looks exactly as before, but you can search it, select it and copy it. Only straightening crooked pages (off unless you choose it) and “Read every page again” save the page pictures again.

    Which languages can be read?

    English, German, French, Spanish, Italian, Portuguese, Dutch, Polish, Romanian, Czech, Hungarian, Turkish, Swedish, Russian, Ukrainian, Greek, Chinese (simplified and traditional), Japanese, Korean, Arabic and Hindi. Choose a second language for documents that mix two.

    How accurate is it?

    Clean printed or typed pages scanned at 300 dpi are read almost word for word. Handwriting, very small print, photos taken at an angle, faded carbon copies and unusual fonts give more mistakes, so check names and numbers before you rely on them.

    Can I OCR a password-protected PDF?

    No. Passwords are not asked for, so a PDF that needs one to open, or whose author set a permissions password against changes, is not changed. Remove the protection in a PDF program with the password, then upload the unprotected copy.

    How many pages can I upload?

    Up to 1,000 pages in one job. A typical 300 dpi page takes one to two seconds to read, so a 100-page scan is done in about three minutes.

    What happens to my files?

    They are read on our server and deleted: the uploads as soon as the job finishes, the results after 30 minutes (or at once with “Delete now”). They are not used for anything else, and the text is not kept.