Skip to content
OCR2text

PDF to text

Get the text out of a PDF — scanned or not

OCR2text opens the PDF in your browser, works out whether it already contains real text, and only runs recognition when it actually needs to. Nothing is uploaded, however many pages it has.

How it goes

  1. 1

    Add the PDF

    Opened locally with PDF.js. Page thumbnails appear straight away.

  2. 2

    Choose pages

    All of them, or a range like 3–7. Big documents are flagged first.

  3. 3

    Text or OCR

    If real text is already there, extract it — faster and perfectly accurate.

  4. 4

    Recognise

    Scanned pages are rasterised and recognised page by page.

  5. 5

    Export

    One combined document, or a searchable PDF with a text layer added.

Text PDFs and scanned PDFs are not the same problem

A PDF exported from Word already contains the characters — running OCR over it would throw away a perfect copy and replace it with a guess. A PDF from an office scanner contains photographs of pages and nothing else, so recognition is the only way in.

OCR2text samples several pages when the file opens and tells you which one you have. If it finds real text you get a choice: Extract existing text, or Run OCR anyway when the text layer is garbled — as it sometimes is on documents that were OCR’d badly years ago.

Large documents

Rasterising and recognising a 200-page scan in a browser tab is real work. Rather than imposing an arbitrary small limit, OCR2text estimates the job from your device’s memory and processor count, warns you when it will be heavy, and suggests a page range.

Pages are processed through a controlled queue — one at a time on a phone, a small pool on a desktop — so the tab does not fall over halfway through.

Searchable PDF

Under Advanced export, OCR2text can build a PDF that looks exactly like the original but carries an invisible text layer, so it becomes searchable and selectable. It re-runs recognition to do this, and it is assembled in your browser like everything else.

Questions

Are my PDF pages sent to a server?+
No. PDF.js parses and renders the file locally, and recognition runs in the same tab. The network panel will show no upload.
Can it open password-protected PDFs?+
Encrypted PDFs cannot be opened. Remove the password in your PDF reader first, then add the file.
Why does a 200-page PDF take so long?+
Because your device is doing the work rather than a rack of servers. That is the trade for not uploading the document. Selecting a page range is usually the right move.

Extract text from something

No account, no upload, no waiting for a server. Open the workspace and drop a page in.

Choose a PDF