Multi-modal OCR pipeline optimized for ML training (text, figure, math, tables, diagrams)
Set up a system to extract text, figures, and tables from documents for educational or research purposes.
Claim its page: indexed whatever its rank, translated into six languages, and enriched with what you write yourself.
Get an email alert on its next release or when it starts trending — never miss the moment.
Free · no card · unsubscribe anytimeMulti-modal OCR pipeline optimized for ML training (text, figure, math, tables, diagrams)
Versatile-OCR-Program has 677 stars on GitHub. It has been forked 50 times. Versatile-OCR-Program is written mainly in Python. It has been in active development since 2025. Its main topics are doclayout, educational-data, exam-ocr, machine-learning.
Multi-modal OCR pipeline optimized for ML training (text, figure, math, tables, diagrams)
Versatile-OCR-Program is an open-source project.
Yes. Versatile-OCR-Program is free and open source — you can use, modify and self-host it.
Versatile-OCR-Program is written mainly in Python.
Add this live badge to your README — your GitHub stars and directory rank, refreshed daily.
[](https://olud.ai/project/raphael-seo-versatile-ocr-program.html)