A scanned PDF looks like a document but behaves like a photograph. You cannot search it, copy a quotation from it, or feed it to an AI summariser, because as far as software is concerned each page is just pixels. Optical Character Recognition fixes that by identifying the shapes of letters and writing a real text layer behind the image.
PDFalot's OCR PDF keeps the original scan visible exactly as it is and places recognised, invisible text underneath at the same coordinates. The page looks unchanged, but Ctrl+F now works, text is selectable, and every AI tool on the site can read the document.
How to use OCR PDF
- 1Upload the scanned PDF or the images you have already converted to PDF.
- 2Select the document's primary language so the recogniser loads the right character model.
- 3Start recognition and watch the per-page progress — OCR is the most computationally heavy tool here.
- 4Download the searchable PDF and test it by searching for a word you can see on the page.
How recognition accuracy depends on your scan
OCR quality is set long before the file reaches this tool. A flat, well-lit 300 DPI scan of printed text routinely exceeds 98% character accuracy. A skewed phone photo of a faint fax at 150 DPI may fall below 80%. If accuracy matters, rescan at 300 DPI in greyscale, keep the page flat, and avoid shadows.
Handwriting and unusual layouts
The engine is trained on printed type. Cursive handwriting is not reliably recognised. Multi-column layouts, tables and forms are read, but reading order in complex layouts can differ from the visual order, so check any extracted table data before relying on it.
OCR unlocks the rest of the toolkit
Once a scan has a text layer, AI Smart Split can find document boundaries, the AI Summarizer can read it, Smart Translate can translate it, and Redact PDF can find and destroy specific words. Running OCR first is often the step that makes everything else work.
Supported formats, limits and things to know
- Processing takes roughly one to several seconds per page depending on your device.
- Handwritten content is not reliably recognised.
- Very low-resolution scans (below about 150 DPI) produce poor results.
- The recognised layer is a best-effort transcription — always verify critical figures against the image.
OCR PDF — frequently asked questions
What does OCR do to a PDF?
It reads the letter shapes in a scanned image and writes them as a hidden, selectable text layer positioned behind the picture, making the document searchable without changing how it looks.
Why is my scanned PDF not searchable?
Because it contains only images of text, not text. Running OCR adds the missing text layer and makes search work.
Which languages are supported?
The recogniser covers the major Latin-script languages plus a wide set of additional character sets. Choose the closest match to your document before starting.
Will OCR change how my document looks?
No. The original scanned image stays exactly as it is on top; the recognised text sits invisibly behind it.
How accurate is it?
A clean 300 DPI scan of printed text typically exceeds 98% character accuracy. Poor scans, faint print and handwriting are considerably lower.
Does my scan get uploaded?
No. Recognition runs in your browser, which is why processing speed depends on your device rather than our servers.