Chandra — распознавание текста и структуры документов

★ 12.1K

Chandra is an open-source text-recognition model that turns images and PDFs into structured text with page layout preserved. Reach for it when a scan or a picture must yield not just letters but structure — headings, tables, column order — so the result is usable downstream without manual reassembly. It returns Markdown, HTML or JSON. It supports over 90 languages and handles complex layouts confidently where ordinary recognition loses tables and confuses columns. It helps digitise contracts, reports, forms and scientific papers and prepare documents for knowledge bases and search. It is a model to embed in a processing pipeline; for a one-off manual cleanup of a single scan, simpler tools fit.