Zum Hauptinhalt springen
Text recognition handles printed pages well, messy handwriting less so - proofread the extracted text before you rely on it. Runs locally on your device.
PlusPDF Tools

PDF Text Recognition

The tax notice is scanned, but you cannot select the numbers, copy them or search them? PDF Text Recognition reads the text out of your scanned PDF - page by page, with previews, confidence scores and full-text search, tuned for German government mail.

A live look inside

Live preview. It becomes interactive with your account.

What PDF Text Recognition does

A scanned PDF is basically a photo. To the eye it looks like text, but the computer only sees pixels - you cannot select anything, copy anything or search anything. OCR (optical character recognition) flips that around: it recognises the letters in the image and turns them into real, searchable text. That is exactly what PDF Text Recognition does, page by page and with previews so you can see what was recognised.

The focus is German government and finance mail. Tax notices, letters from the authorities, contracts, certificates, invoices - all of it often arrives scanned and is exactly then most annoying, when you want to reuse the numbers and deadlines inside. Ready-made presets set language and resolution to fit the document type so you do not have to fiddle with settings.

Recognition runs in German, English or a mix of both. German texts with umlauts and the sharp ß in particular benefit from choosing the right language. For documents that mix both languages there is the combined mode, which uses both dictionaries at once.

Resolution decides accuracy and speed. 150 DPI is fast and enough for clean, clearly printed pages. 300 DPI is the standard for most documents. 600 DPI squeezes out the maximum from small print, contracts or slightly blurry scans, but takes a bit longer. Every page comes with a confidence score that shows you how sure the recognition was.

You can view the result page by page or as one continuous text, search it right in the browser and pull the recognised text out - ready to reuse, say to copy a figure into a calculator or move a deadline into the Fristenwächter. Word and character counts per page are always in view.

The actual character recognition needs computing power and therefore runs on the server, but only for the pure recognition step. Your documents are not stored permanently. No subscription for a simple scan, no detour through a foreign cloud service with murky privacy terms.

Features

Per-page recognition

Each page is recognised individually and shown with a thumbnail so you can compare text and original.

German, English or both

Choose the language to match the document. Combined mode uses both dictionaries at once.

Presets for official mail

Government letter, tax notice, contract or invoice - each with a fitting language and resolution.

Three resolution levels

150 DPI fast, 300 DPI standard, 600 DPI for small print and difficult scans.

Confidence score per page

A value per page shows how confident the recognition was - so you know where to double-check.

Full-text search

Search the recognised text right in the browser and find a number, a name or a deadline instantly.

Pull the text out

View the result per page or as one block and take the text to reuse elsewhere.

How it works

  1. 1

    Load a scanned PDF

    Pick the scanned document. Even multi-page scans up to 50 MB are no problem.

  2. 2

    Choose language and quality

    German, English or both, plus the resolution. A preset like tax notice sets both to fit.

  3. 3

    Start recognition

    The text is recognised page by page. For each page you see preview, confidence and text volume.

  4. 4

    Search and reuse

    Search the text, check uncertain spots and take the result to reuse elsewhere.

Who needs this

→Anyone reusing figures from a scanned tax notice.
→People wanting to make a government letter searchable.
→Bookkeepers pulling text out of scanned invoices.
→Anyone making contracts and certificates from the scan archive readable again.
→People digitising paper mail and turning it into real text.

Frequently asked questions

What is OCR exactly?

OCR stands for optical character recognition. It recognises letters in an image or scanned PDF and turns them into real, searchable and copyable text. A photo of writing becomes writing the computer can work with again.

Does it recognise German umlauts?

Yes. Choose German as the language and ä, ö, ü and ß are recognised correctly. For documents mixing German and English there is the combined mode.

Which resolution should I use?

300 DPI is the standard and fits most documents. Use 150 DPI for cleanly printed pages when speed matters, and 600 DPI for small print, contracts or slightly blurry scans.

What does the confidence score mean?

The confidence score shows how sure the recognition was for a page. A low value is a hint to check that page more closely, for example with a poor scan or unusual typeface.

Do my documents stay private?

The pure character recognition needs computing power and runs on the server, but only for that step. Your documents are not stored permanently and are not passed on to third-party services.

Can I process multi-page scans?

Yes. The tool processes multi-page PDFs up to 50 MB and shows the result for each page individually as well as one continuous text.

Image Text OCR

Recognize and extract text from images, scans, and screenshots. German, English, or mixed. Se…

PDF to Image

Convert PDF pages to high-resolution JPG, PNG, or WebP images. Choose DPI and page range. Fre…

Document Scanner

Photograph documents with your phone, mark corners, deskew server-side. Batch scanning, prese…

PDF Editor

Edit PDF files directly in your browser: text, highlights, underlines, strikethroughs, rectan…

Ready to use PDF Text Recognition?

No installation. No account needed to start. Open it right in your browser.

Open now