Local embedded-text workspace
Extract text from PDF
Extract embedded text from a PDF.
- Processing
- Browser only
- Account
- Not required
- Current guardrail
- One PDF · 100 MB · 250 pages
- Output
- A UTF-8 text file
No file upload, analytics, ads or external processing scripts in this workspace.
Local embedded-text extraction
Choose one PDF
PigPDF reads the embedded text layer in source page order and creates a separate UTF-8 text file.
or choose a file from this device
Current compatibility
What this beta accepts
PigPDF reads the embedded text stream without OCR. The output is plain text, not a reconstruction of the page layout.
| Input feature | Current behavior | Reason |
|---|---|---|
| Embedded page text | Extracted | Selected pages are read in source page order. |
| Image-only scanned pages | No text | OCR is a separate planned capability. |
| Complex columns and visual layout | May change | PDF stream order is not guaranteed to match semantic reading order. |
| Password protection or damage | Rejected | No password entry or repair path is included in this beta. |
Before you use this tool
Questions and limits
Does my file leave this device?
No file content is sent by this tool. The active processing path runs in this browser and is covered by an automated egress test.
Does PigPDF change my source file?
No. The source stays unchanged and PigPDF creates a separate output file.
Which current limits apply?
One PDF up to 100 MB and 250 pages. Only embedded text is extracted; scans require OCR and complex reading order may be ambiguous.