OCR & Scanned Docs: Searchable PDFs & Text Layers
Optical Character Recognition, text extraction, and scanned document searchability.
Files are processed locally in your browser with WebAssembly and HTML5 APIs. They never leave your device, so there are no server uploads or third-party file storage.
Start converting immediately on your own device - no upload wait, no download queue, and no throttling by a remote server.
Use every available tool without creating an account, paying for a plan, or hitting an artificial file-size limit. Convert as often as your browser and device can handle.
Format Comparison & Technical Specifications
| Output Format | MIME Type | Specification | Type |
|---|---|---|---|
| PDF Scanned | application/pdf | ISO-32000-2 | Bitmap Scanned PDF |
| IMG Scanned Image | image/jpeg | W3C-PNG | Scanned Document Image |
| TXT Extracted Text | text/plain | RFC-4180-CSV | Extracted Text Output |
| DOCX Word Document | application/vnd.openxmlformats-officedocument.wordprocessingml.document | ISO-IEC-29500 | Structured Word Document |
| HOCR HTML | text/html | HOCR-SPEC-12 | hOCR Spatial Markup |
| PNG Document | image/png | W3C-PNG | High-Resolution Scanned PNG |
| JPG Document | image/jpeg | W3C-PNG | Scanned Document JPEG |
| PDF Searchable | application/pdf | ISO-32000-2 | OCR Searchable Document |
Technical Specifications & Codec Breakdown
How Searchable PDFs Work
A scanned document is just a picture of paper. An OCR engine recognizes letter shapes and inserts an invisible, selectable text layer right on top of the image. When you highlight or search text in the PDF, you are selecting that invisible text layer.
The 300 DPI Sweet Spot for Scanning
Scanning documents at 150 DPI causes broken letters and poor OCR accuracy. Scanning at 600 DPI quadruples file size without improving text recognition. 300 DPI is the universal sweet spot for high OCR accuracy and compact file sizes.
0 bytes uploaded to external servers. All processing occurred locally in your browser sandbox.