pip install --upgrade pip pip install uv uv pip install -U "mineru[all]"
Excerpts from the project README on GitHub. Copyright and licensing remain with the respective authors.
MinerU extracts text, tables, formulas and images from PDFs into Markdown or JSON, handling scientific documents particularly well.
| Category | Document AI & OCR |
| Type | Document extractor |
| License | AGPL-3.0 |
| Runs locally | Yes |
| Built with | Python |
| Skill level | Intermediate |
| Best for | scientific PDFs with formulas |
Other open-source document ai & ocr tools worth comparing:
MarkerConvert PDFs to clean Markdown fast
DoclingIBM-grade document understanding
SuryaModern OCR with layout detection
PaddleOCRIndustrial OCR with tiny models
UnstructuredTurn any file into LLM-ready data
FirecrawlTurn websites into LLM-ready Markdown
Crawl4AIOpen-source LLM-friendly crawler
olmOCRTurn PDFs into clean training-grade text
TesseractThe classic OCR engineMinerU is free and open-source (AGPL-3.0 license), so you can use, self-host and modify it at no cost.
Yes. MinerU is designed to run on your own machine or server, keeping your data private.
Popular open-source alternatives include Marker, Docling, Surya. See the comparisons above to choose.
Browse the full directory of open-source AI tools, models and projects — updated daily.
Browse all tools →