Home Projects PaddleOCR
PaddleOCR

PaddleOCR

by PaddlePaddle · GitHub

Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.

# ai4science# chineseocr# document-parsing# document-translation
View on GitHub
⭐ Stars
86k
🍴 Forks
11.1k
🔥 Trending
+490 this week
📜 License
Apache-2.0
Commercial use OK
📅 Created
2020
🔄 Last commit
6 days ago
🏷️ Category
ai4science
💻 Language
You maintain this project?

Claim its page: indexed whatever its rank, translated into six languages, and enriched with what you write yourself.

Claim this page →
PaddleOCR — GitHub preview card
📈 Star history
86k84k
2026-06-272026-07-22
📄 About

Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.

Read the full guide
Frequently asked questions

What is PaddleOCR?

Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.

Is PaddleOCR open source?

PaddleOCR is an open-source project. It is released under the Apache-2.0 license.

Is PaddleOCR free?

Yes. PaddleOCR is free and open source — you can use, modify and self-host it.

🏅 Maintainer of this project?
OpenSourceAI badge — PaddleOCR

Add this live badge to your README — your GitHub stars and directory rank, refreshed daily.

[![OpenSourceAI](https://opensourceai.tech/badge.php?tool=paddlepaddle-paddleocr)](https://opensourceai.tech/project/paddlepaddle-paddleocr.html)
More badge options →
🧬 Related projects🧬 View the DNA map →