Turn scanned newspaper PDFs into structured text.
Clone the repository and install:
pip install -e .Make an input directory:
mkdir inputCopy your newspaper .pdf file(s) into it, either through using a file manager or through your CLI:
cp ~/path/to/newspapers/*.pdf input/Batch process all PDFs in input/:
compositor batch
# alternatively, if the above doesn't work:
python3 -m compositor batchResults land in output/<pdf-stem>/, where <pdf-stem> is the PDF filename.
| command | what it does |
|---|---|
process <pdf> |
Process a single PDF |
batch |
Process every PDF in input/ (--force to redo) |
rebuild <issue_id> |
Reassemble documents from cached layout.json, no re-OCR |
group |
List the corpus grouped by publication & date |
stats |
Summarize the index.json manifest |
export |
Write corpus/corpus.jsonl (--use-corrected, --min-word-count N) |
clean <issue_id> |
Optional LLM OCR correction via a local Ollama server |
Built on Surya which does the layout analysis and text recognition this pipeline builds its document structure from. Also uses PyMuPDF for rendering and embedded-text extraction.