OCRmyPDF — распознавание текста в сканах PDF

★ 34.5K

OCRmyPDF is a tool that adds a recognized text layer (OCR) to scanned PDFs, making them searchable and copyable. Reach for it when you have scans or paper documents captured as PDF that you can't search: OCRmyPDF runs them through recognition and returns the same PDF but with an invisible text layer over the image — nothing changes visually, yet you can now search, select, and copy the text. It works from the command line, slots easily into pipelines, uses the Tesseract engine, and can clean and deskew scans. Its focus is turning image-only PDFs into searchable ones while preserving appearance, not extracting tables (excalibur) or layout analysis/chunking for RAG (docling/chunkr): a narrow, reliable "scan → searchable PDF" job. It is often paired with systems like paperless-ngx to automatically process incoming documents.