olmOCR from AI2 converts messy PDFs into clean linearised text at scale, preserving reading order across columns and figures.
pip install olmocr
Excerpts from the project README on GitHub. Copyright and licensing remain with the respective authors.
Claim its page: indexed whatever its rank, translated into six languages, and enriched with what you write yourself.
Get an email alert on its next release or when it starts trending — never miss the moment.
Free · no card · unsubscribe anytimeToolkit for linearizing PDFs for LLM datasets/training
olmocr has 19.4k stars on GitHub. It has been forked 1.6k times. olmocr is written mainly in Python. It has been in active development since 2024. olmocr is available under the Apache-2.0 license.
Read the full guideToolkit for linearizing PDFs for LLM datasets/training
olmocr is an open-source project. It is released under the Apache-2.0 license.
Yes. olmocr is free and open source — you can use, modify and self-host it.
olmocr is available under the Apache-2.0 license.
olmocr is written mainly in Python.
Add this live badge to your README — your GitHub stars and directory rank, refreshed daily.
[](https://olud.ai/project/allenai-olmocr.html)