Local OCR workspace
Extract Scanned Text
Run the bundled local English OCR model and export recognized page text.
- Processing
- Browser only
- Account
- Not required
- Current guardrail
- PDF · 10 pages · 60 MP · English model
- Output
- A UTF-8 OCR text file
No file upload, analytics, ads or external processing scripts in this workspace.
Local browser workspace
Choose source files
Run this operation in the browser and create a separate result file.
or choose files from this device
Local browser workspace
The source files are not changed and are not uploaded.Current compatibility
What this beta accepts
Pages are rasterized locally and recognized by an English OCR model bundled with PigPDF.
| Input feature | Current behavior | Reason |
|---|---|---|
| PDF pages | Rasterized and recognized | PDF.js renders each page locally before the OCR worker reads it. |
| Printed English text | Recognized | The bundled English model produces text and word positions. |
| Handwriting, other languages or complex layout | Review required | Accuracy varies and no language model is silently selected. |
| Forms, annotations and signatures | Rasterized | The searchable output is a new image-based PDF and does not preserve interactive semantics. |
Before you use this tool
Questions and limits
Does my file leave this device?
No file content is sent by this tool. The active processing path runs in this browser and is covered by an automated egress test.
Does PigPDF change my source file?
No. The source stays unchanged and PigPDF creates a separate output file.
Which current limits apply?
One PDF up to 50 MB and 10 pages. The UTF-8 output is recognized by the bundled English model and is not a certified transcription or reading order.