Scanned court filings — thousands of pages, no text layer, and material too sensitive to send to a hosted model.
A local-only pipeline: OCR, layout-aware chunking, embeddings and retrieval that never leave the machine, with citations back to the page image.
10,062 pages of scanned petitions turned into structured text — 2.8× what the PDFs' own text layer carried.