pip install docling

Excerpts from the project README on GitHub. Copyright and licensing remain with the respective authors.
Docling parses PDFs, DOCX, PPTX and images into structured representations with layout, tables and reading order, designed to plug straight into RAG frameworks.
| Category | Document AI & OCR |
| Type | Document parsing toolkit |
| License | MIT |
| Runs locally | Yes |
| Built with | Python |
| Skill level | Intermediate |
| Best for | teams building document pipelines for RAG |
Other open-source document ai & ocr tools worth comparing:
MarkerConvert PDFs to clean Markdown fast
SuryaModern OCR with layout detection
PaddleOCRIndustrial OCR with tiny models
UnstructuredTurn any file into LLM-ready data
FirecrawlTurn websites into LLM-ready Markdown
Crawl4AIOpen-source LLM-friendly crawler
olmOCRTurn PDFs into clean training-grade text
MinerUPDF to Markdown with formulas and tables
TesseractThe classic OCR engineDocling is free and open-source (MIT license), so you can use, self-host and modify it at no cost.
Yes. Docling is designed to run on your own machine or server, keeping your data private.
Popular open-source alternatives include Marker, Surya, PaddleOCR. See the comparisons above to choose.
Browse the full directory of open-source AI tools, models and projects — updated daily.
Browse all tools →