Latest article
pdf-inspector Architecture: Detection, Layout, Tables, Markdown, and OCR
Follow one PDF through single-load parsing, content streams, fonts, layout, tables, Markdown, and selective OCR.
A hands-on series on PDF classification, layout extraction, and selective OCR.

Latest article
Follow one PDF through single-load parsing, content streams, fonts, layout, tables, Markdown, and selective OCR.
By publication date
01 → 09
Follow one PDF through single-load parsing, content streams, fonts, layout, tables, Markdown, and selective OCR.
Choose a PDF ingestion path by layout quality, OCR need, privacy, latency, cost, and operational burden.
Choose local CLI, service, browser WebAssembly, or native bindings while keeping PDFium and ONNX optional.
Run detection, Markdown extraction, table handling, and selective OCR through Python, Node, Rust, or CLI.
Design a local document-ingestion project with classification manifests, selective OCR, provenance, and accessible outputs.
A source-backed overview of Rust parsing, layout extraction, selective OCR, bindings, and benchmarks.
Reproduce corpus speed and quality while separating native parsing, OCR, memory, storage, and hosted fallbacks.
Protect local documents, OCR runtimes, hosted fallbacks, output provenance, and service boundaries.
Use fixtures to read detector, extractor, font decoding, layout, tables, bindings, and OCR boundaries.