OCR & Scanned Docs: Searchable PDFs & Text Layers
Optical Character Recognition, text extraction, and scanned document searchability.
Your files are processed entirely in your web browser using HTML5 Canvas and WebAssembly. Zero data ever leaves your device.
Skip upload and download queues. Conversions happen locally at full machine hardware speed without server limits.
Compliant with official ISO, W3C, and IETF specifications to guarantee output compatibility across modern operating systems.
Format Comparison & Technical Specifications
| Output Format | MIME Type | Specification | Type |
|---|---|---|---|
| PDF Scanned | application/pdf | ISO-32000-2 | Bitmap Scanned PDF |
| IMG Scanned Image | image/jpeg | W3C-PNG | Scanned Document Image |
| TXT Extracted Text | text/plain | RFC-4180-CSV | Extracted Text Output |
| DOCX Word Document | application/vnd.openxmlformats-officedocument.wordprocessingml.document | ISO-IEC-29500 | Structured Word Document |
| HOCR HTML | text/html | HOCR-SPEC-12 | hOCR Spatial Markup |
| PNG Document | image/png | W3C-PNG | High-Resolution Scanned PNG |
| JPG Document | image/jpeg | W3C-PNG | Scanned Document JPEG |
| PDF Searchable | application/pdf | ISO-32000-2 | OCR Searchable Document |
Technical Specifications & Codec Breakdown
How Searchable PDFs Work
A scanned document is just a picture of paper. An OCR engine recognizes letter shapes and inserts an invisible, selectable text layer right on top of the image. When you highlight or search text in the PDF, you are selecting that invisible text layer.
The 300 DPI Sweet Spot for Scanning
Scanning documents at 150 DPI causes broken letters and poor OCR accuracy. Scanning at 600 DPI quadruples file size without improving text recognition. 300 DPI is the universal sweet spot for high OCR accuracy and compact file sizes.
0 bytes uploaded to external servers. All processing occurred locally in your browser sandbox.