Ranking Scanned PDFs While Preserving OCR Uncertainty
Combine lexical and semantic retrieval, then use region evidence and page metadata to rank results without concealing weak extraction.
A scanned PDF presents two separate search problems: finding the
sourcebento.hashnode.dev9 min read