Home Projects olmocr
olmocr
Python

olmocr

Toolkit for linearizing PDFs for LLM datasets/training

by allenai · GitHub
Stars
Forks
License
Created
Last commit
Language
Apache-2.0Python
View on GitHub
In plain words

olmOCR from AI2 converts messy PDFs into clean linearised text at scale, preserving reading order across columns and figures.

From the README

Install
pip install olmocr

Excerpts from the project README on GitHub. Copyright and licensing remain with the respective authors.

You maintain this project?

Claim its page: indexed whatever its rank, translated into six languages, and enriched with what you write yourself.

Claim this page →
olmocr — GitHub preview card
📈 Star history
19.4k19.1k
2026-07-122026-08-31
📈 Track olmocr

Get an email alert on its next release or when it starts trending — never miss the moment.

Free · no card · unsubscribe anytime
Get email alerts →
📄 About

Toolkit for linearizing PDFs for LLM datasets/training

olmocr has 19.4k stars on GitHub. It has been forked 1.6k times. olmocr is written mainly in Python. It has been in active development since 2024. olmocr is available under the Apache-2.0 license.

Read the full guide
Frequently asked questions

What is olmocr?

Toolkit for linearizing PDFs for LLM datasets/training

Is olmocr open source?

olmocr is an open-source project. It is released under the Apache-2.0 license.

Is olmocr free?

Yes. olmocr is free and open source — you can use, modify and self-host it.

What license does olmocr use?

olmocr is available under the Apache-2.0 license.

What language is olmocr written in?

olmocr is written mainly in Python.

🏅 Maintainer of this project?
olud.ai badge — olmocr

Add this live badge to your README — your GitHub stars and directory rank, refreshed daily.

[![olud.ai](https://olud.ai/badge.php?tool=allenai-olmocr)](https://olud.ai/project/allenai-olmocr.html)
More badge options →
🧬 Related projects