OCR PDF
OCR PDF to searchable text — recognize scanned documents with Tesseract.
Use Cases
Typical scenarios for OCR PDF — see if it fits before you start
OCR PDF · Scenario 1
Turning scans into searchable text, OCR PDF checks text layer then falls back to local Tesseract with auto lang.
OCR PDF · Scenario 2
Batch papers/archives, OCR PDF queues with progress and exports to Word/Excel/text.
OCR PDF · Scenario 3
For scanned tables, OCR PDF preserves rows/columns and prompts proofing numbers.
How to Use
- Upload — Click or drag PDF/images to the dashed zone. Multi-file and large files (<100MB per batch recommended) are read locally only.
- Configure & process — Set pages, angle, password or quality, then hit “Calculate / Process”. OCR PDF runs via WebAssembly and pdf-lib in your browser — no server queue.
- Preview & download — Check pages, size and thumbnails on the right, then download. Batch jobs download one by one; tweak and re-run if needed.
Tips & Notes
- Page syntax: Use
1,3,5-7or1-10,15; empty means all. Out-of-range pages are ignored. - Privacy & size: Everything runs locally — no server upload. For large files, process in batches of 50–100 pages.
- Fidelity: Fonts and vectors are preserved where possible. For scanned PDFs, run OCR/searchable PDF first, then convert.
機能
- OCR searchable
- Scanned text extraction
- Auto language detect
Supported Formats
All conversions run locally in the browser; check fidelity in desktop for complex layouts.
よくある質問
Scanned not searchable, how to make searchable?
OCR PDF checks text layer, if missing runs local Tesseract (on-demand tesseract.min.js) with auto language, generating searchable layer — for contracts and invoices. Preview pages first, then by “1-20” to avoid memory spikes.
Keep tables and layout?
Best effort — paragraphs and basic table rows/columns restored, complex headers better via Word/Excel export — for reports. Preview after OCR, proofread key numbers before distribution.
Batch & offline?
Yes. Multi-page queue with progress, fully offline, no upload — for sensitive archives. Batch 20 pages at a time, close tab to clear memory per audit.
使い方
OCR PDF to searchable text — recognize scanned documents with Tesseract.
すべての計算はブラウザ内で実行され、データはアップロードされません。