Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
MinerU extracts text, tables, formulas and images from PDFs into Markdown or JSON, handling scientific documents particularly well.
pip install --upgrade pip pip install uv uv pip install -U "mineru[all]"
Excerpts from the project README on GitHub. Copyright and licensing remain with the respective authors.
Claim its page: indexed whatever its rank, translated into six languages, and enriched with what you write yourself.
Get an email alert on its next release or when it starts trending — never miss the moment.
Free · no card · unsubscribe anytimeTransforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
MinerU has 78.8k stars on GitHub. It has been forked 6.6k times. MinerU is written mainly in Python. It has been in active development since 2024. Its main topics are ai4science, document-analysis, docx, extract-data.
Read the full guideTransforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
MinerU is an open-source project.
Yes. MinerU is free and open source — you can use, modify and self-host it.
MinerU is written mainly in Python.
Add this live badge to your README — your GitHub stars and directory rank, refreshed daily.
[](https://olud.ai/project/opendatalab-mineru.html)
Measured from GitHub topics shared by both projects, weighted by how rare each topic is.