German OCR: Read German Scans With Umlauts (Beta)

Read German letters, forms and old typed documents, with ä, ö, ü and ß recognised correctly.

Opens: PDF, JPG, PNG, TIFF, WebP, BMP, HEIC, AVIF
Saves: PDF, TXT

Read the text

Scanned PDFs or pictures
Choose files… or drop them here
.pdf, .jpg, .jpeg, .jpe, .jfif, .png, .tif, .tiff, .webp, .bmp, .heic, .heif, .hif, .avif · up to 200 MB each, 200 files max
    Add more files

    PDFs are read page by page and each gives a searchable PDF. Pictures (photos, screenshots, scans; several or a multi-page TIFF) become one PDF, in the order you add them.

    Options

    Choose the language the text is written in: letters with accents and whole words are read much better.

    For documents that mix two languages, such as English terms in a Russian letter. Reading takes about a quarter longer.

    Each page's text direction is checked, and pages scanned the wrong way round are turned. The page pictures themselves are not changed.

    Pages scanned at a slight angle are straightened. Their pictures are then saved again, so the PDF can get larger.

    Typed PDFs and pages read before have text you can select already.

    More options

    Leave empty for all pages, or write pages like 1-3, 5, 8- (8 to the end). The other pages stay in the PDF as they are.

    Only for pictures: how they become PDFs.

    Your files are deleted after processing.

    German text read with an English OCR model loses its umlauts and German quotation marks: in our test scan, "über den faulen Hund" came out as "iiber den faulen Hund" and „schnelle” as ,.schnelle”. Reading with the German language model fixes that: ä, ö, ü and ß are expected, and the German dictionary helps with long words.

    How to OCR a German document

    1. Add the scanned PDF or pictures above. The language is already set to German.
    2. For a document that also has English (or French) passages, add them as the second language.
    3. Press Read the text and download the searchable PDF and the text.

    What it handles

    • Umlauts and ß in modern typefaces, on letters, invoices, contracts and forms.
    • Typewritten documents from the typewriter era, when the scan is clean.
    • Mixed documents, such as a German letter with an English product name, with a second language.

    Limits

    Blackletter printing (Fraktur), as in books and newspapers from before the 1940s, is a different script with its own letter shapes and is read poorly with the modern German model. Handwriting (including Kurrent and Sütterlin) can't be read.

    Tips

    • Scan at 300 dpi in grey or black and white; colour adds nothing for text.
    • Straighten crooked pages under the options if the lines run downhill.
    • Check umlauts in names and places against the scan, especially in small print: dots are the first thing a poor scan loses.

    More about what this tool reads and writes: OCR PDF.