pip install olmocr
Excerpts from the project README on GitHub. Copyright and licensing remain with the respective authors.
olmOCR from AI2 converts messy PDFs into clean linearised text at scale, preserving reading order across columns and figures.
| Category | Document AI & OCR |
| Type | OCR pipeline |
| License | Apache-2.0 |
| Runs locally | Yes |
| Built with | Python |
| Skill level | Intermediate |
| Best for | bulk PDF extraction for RAG or training |
Other open-source document ai & ocr tools worth comparing:
MarkerConvert PDFs to clean Markdown fast
DoclingIBM-grade document understanding
SuryaModern OCR with layout detection
PaddleOCRIndustrial OCR with tiny models
UnstructuredTurn any file into LLM-ready data
FirecrawlTurn websites into LLM-ready Markdown
Crawl4AIOpen-source LLM-friendly crawler
MinerUPDF to Markdown with formulas and tables
TesseractThe classic OCR engineolmOCR is free and open-source (Apache-2.0 license), so you can use, self-host and modify it at no cost.
Yes. olmOCR is designed to run on your own machine or server, keeping your data private.
Popular open-source alternatives include Marker, Docling, Surya. See the comparisons above to choose.
Browse the full directory of open-source AI tools, models and projects — updated daily.
Browse all tools →