Skip to main content
Document Analysis — Overview
Purpose: Turn heterogeneous document drops (FOIA, leaks, archives) into searchable, citable evidence.
Core pipeline
- Inventory — List formats (PDF, email, images, spreadsheets).
- Extract text — Born-digital PDFs vs. scans (OCR).
- Normalize — Consistent filenames, date fields, deduplication keys.
- Index — Full-text index; optional entity extraction downstream.
- Search — Iterative queries; save search strings in a research log.
When to escalate
| Situation | Action |
|---|---|
| >100k pages, team collaboration | Consider Datashare, Aleph, or dedicated investigation platform |
| Audio/video central to case | Plan transcription pass; see Transcription Services |
| Messy relational data | OpenRefine before joining to corporate or FEC tables |
Related skill
document-research-specialist
