BETA

Local embedded-text workspace

Extract text from PDF

Extract embedded text from a PDF.

Processing
Browser only
Account
Not required
Current guardrail
One PDF · 100 MB · 250 pages
Output
A UTF-8 text file
Browser-only processing

No file upload, analytics, ads or external processing scripts in this workspace.

IDLE

Local embedded-text extraction

Choose one PDF

PigPDF reads the embedded text layer in source page order and creates a separate UTF-8 text file.

Drop one PDF here

or choose a file from this device

One PDF up to 100 MB and 250 pages · scanned pages require OCR

Current compatibility

What this beta accepts

PigPDF reads the embedded text stream without OCR. The output is plain text, not a reconstruction of the page layout.

PDF to TextCurrent compatibility
Input featureCurrent behaviorReason
Embedded page textExtractedSelected pages are read in source page order.
Image-only scanned pagesNo textOCR is a separate planned capability.
Complex columns and visual layoutMay changePDF stream order is not guaranteed to match semantic reading order.
Password protection or damageRejectedNo password entry or repair path is included in this beta.

Before you use this tool

Questions and limits

Does my file leave this device?

No file content is sent by this tool. The active processing path runs in this browser and is covered by an automated egress test.

Does PigPDF change my source file?

No. The source stays unchanged and PigPDF creates a separate output file.

Which current limits apply?

One PDF up to 100 MB and 250 pages. Only embedded text is extracted; scans require OCR and complex reading order may be ambiguous.