Skip to main content
OCR Tools
| Tool | License / access | Best for |
|---|---|---|
| Tesseract | Open source | Bulk OCR; language packs; local control |
| Google Cloud Vision | Commercial API | Hard scans, mixed layouts |
| Amazon Textract | Commercial API | Forms and tables extraction |
| DocumentCloud (add-ons) | Platform | Hosted investigative docs with OCR pipeline |
Quality checklist
- Sample 10 random pages per scanner batch; measure character error rate subjectively.
- Watch for column order swaps in newspapers and financial tables.
- Store OCR version and language model in the research log.
Provenance
OCR output is machine-derived. Label uncertain readings when quoting in published analysis; prefer image snippet + transcript side-by-side for high-stakes claims.
