Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
PaddleOCR is an industrial-strength multilingual OCR toolkit with ultra-lightweight models that run from servers down to edge devices.

Excerpts from the project README on GitHub. Copyright and licensing remain with the respective authors.
Claim its page: indexed whatever its rank, translated into six languages, and enriched with what you write yourself.
Get an email alert on its next release or when it starts trending — never miss the moment.
Free · no card · unsubscribe anytimeTurn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
PaddleOCR has 88.5k stars on GitHub. It has been forked 11.3k times. PaddleOCR is written mainly in Python. It has been in active development since 2020. PaddleOCR is available under the Apache-2.0 license. Its main topics are ai4science, chineseocr, document-parsing, document-translation.
Read the full guideTurn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
PaddleOCR is an open-source project. It is released under the Apache-2.0 license.
Yes. PaddleOCR is free and open source — you can use, modify and self-host it.
PaddleOCR is available under the Apache-2.0 license.
PaddleOCR is written mainly in Python.
Add this live badge to your README — your GitHub stars and directory rank, refreshed daily.
[](https://olud.ai/project/paddlepaddle-paddleocr.html)
Measured from GitHub topics shared by both projects, weighted by how rare each topic is.