Skip to content

PDF OCR

Read the text out of a scanned PDF.

The first run downloads the OCR engine and your chosen language model — about 12–15 MB — from a public CDN, then your browser caches it. The recognition itself happens on your device: the PDF is never uploaded.

Picking the right language makes a large difference to accuracy.

Drop a scanned PDF

Up to 20 pages per run · nothing is uploaded

Nothing read yet

Add a scanned PDF and its pages are read as text.

Private by design. This tool runs entirely in your browser. Your file is processed on your own device and is not uploaded by DO101.

About the PDF OCR

Runs optical character recognition over each page of a scanned PDF and gives you back copyable text, with a confidence score so you know how much to trust it. Seventeen languages, and the whole thing runs on your device.

How to use it

  1. 1Choose the language of the document.
  2. 2Add the scanned PDF, up to 20 pages per run.
  3. 3Wait for recognition, then copy or download the text.

What you get

  • 17 languages
  • Per-run confidence score
  • Pages rendered at high DPI for accuracy
  • Copy or download as .txt
  • Recognition runs on your device

Frequently asked questions

Why is the first run slow?

The OCR engine and language model — about 12 to 15 MB — are downloaded from a public CDN the first time. Your browser caches them, so later runs start immediately.

How accurate is it?

On a clean, straight, high-contrast scan, very good. On a crooked phone photo of a faint receipt, poor. The confidence score tells you which situation you are in, and DO101 warns you when it drops below 70%.

Why only 20 pages?

OCR is heavy work. Capping the batch keeps the browser responsive and prevents the tab from running out of memory. Split a longer document first.

Is my document uploaded?

No. Only the engine and language model are downloaded; your PDF stays on your device.

People who use this tool usually reach for these next.