pip install marker-pdf
Excerpts from the project README on GitHub. Copyright and licensing remain with the respective authors.
Marker converts PDFs and office documents to clean Markdown, JSON and HTML with high accuracy, handling tables, equations and multi-column layouts.
| Category | Document AI & OCR |
| Type | PDF-to-Markdown converter |
| License | GPL-3.0 |
| Runs locally | Yes |
| Built with | Python |
| Skill level | Intermediate |
| Best for | developers feeding documents into RAG pipelines |
Other open-source document ai & ocr tools worth comparing:
DoclingIBM-grade document understanding
SuryaModern OCR with layout detection
PaddleOCRIndustrial OCR with tiny models
UnstructuredTurn any file into LLM-ready data
FirecrawlTurn websites into LLM-ready Markdown
Crawl4AIOpen-source LLM-friendly crawler
olmOCRTurn PDFs into clean training-grade text
MinerUPDF to Markdown with formulas and tables
TesseractThe classic OCR engineMarker is free and open-source (GPL-3.0 license), so you can use, self-host and modify it at no cost.
Yes. Marker is designed to run on your own machine or server, keeping your data private.
Popular open-source alternatives include Docling, Surya, PaddleOCR. See the comparisons above to choose.
Browse the full directory of open-source AI tools, models and projects — updated daily.
Browse all tools →