Documind — извлечение данных из PDF в JSON

★ 1.5K

Documind is a document-processing tool that uses AI to extract structured data from PDFs. Reach for it when you need to turn unstructured documents into data on your own schema: convert PDFs to Markdown, extract the fields you need, and get structured JSON per a defined (customizable) schema — for example, parse invoices, contracts, and forms for further processing. Its focus is AI extraction of structured data from documents by a custom schema, not a general parser to Markdown (docling) or OCR (OCRmyPDF): the value is getting predictable JSON from documents for your fields. In job it sits alongside Unstract, Sparrow, and chunkr; choose by how the schema is defined and by deployment. It suits people who need to automatically pull fields from PDFs into structured data for their systems.