German text read with an English OCR model loses its umlauts and German quotation marks: in our test scan, "über den faulen Hund" came out as "iiber den faulen Hund" and „schnelle” as ,.schnelle”. Reading with the German language model fixes that: ä, ö, ü and ß are expected, and the German dictionary helps with long words.
How to OCR a German document
- Add the scanned PDF or pictures above. The language is already set to German.
- For a document that also has English (or French) passages, add them as the second language.
- Press Read the text and download the searchable PDF and the text.
What it handles
- Umlauts and ß in modern typefaces, on letters, invoices, contracts and forms.
- Typewritten documents from the typewriter era, when the scan is clean.
- Mixed documents, such as a German letter with an English product name, with a second language.
Limits
Blackletter printing (Fraktur), as in books and newspapers from before the 1940s, is a different script with its own letter shapes and is read poorly with the modern German model. Handwriting (including Kurrent and Sütterlin) can't be read.
Tips
- Scan at 300 dpi in grey or black and white; colour adds nothing for text.
- Straighten crooked pages under the options if the lines run downhill.
- Check umlauts in names and places against the scan, especially in small print: dots are the first thing a poor scan loses.
More about what this tool reads and writes: OCR PDF.