BETA

Local OCR workspace

Extract Scanned Text

Run the bundled local English OCR model and export recognized page text.

Processing
Browser only
Account
Not required
Current guardrail
PDF · 10 pages · 60 MP · English model
Output
A UTF-8 OCR text file
Browser-only processing

No file upload, analytics, ads or external processing scripts in this workspace.

IDLE

Local browser workspace

Choose source files

Run this operation in the browser and create a separate result file.

Drop source files here

or choose files from this device

Up to 100 MB per PDF and 250 pages in this beta

Local browser workspace

The source files are not changed and are not uploaded.

Current compatibility

What this beta accepts

Pages are rasterized locally and recognized by an English OCR model bundled with PigPDF.

Extract Scanned TextCurrent compatibility
Input featureCurrent behaviorReason
PDF pagesRasterized and recognizedPDF.js renders each page locally before the OCR worker reads it.
Printed English textRecognizedThe bundled English model produces text and word positions.
Handwriting, other languages or complex layoutReview requiredAccuracy varies and no language model is silently selected.
Forms, annotations and signaturesRasterizedThe searchable output is a new image-based PDF and does not preserve interactive semantics.

Before you use this tool

Questions and limits

Does my file leave this device?

No file content is sent by this tool. The active processing path runs in this browser and is covered by an automated egress test.

Does PigPDF change my source file?

No. The source stays unchanged and PigPDF creates a separate output file.

Which current limits apply?

One PDF up to 50 MB and 10 pages. The UTF-8 output is recognized by the bundled English model and is not a certified transcription or reading order.