pip install firecrawl-py
Excerpts from the project README on GitHub. Copyright and licensing remain with the respective authors.
Firecrawl crawls websites and returns clean Markdown or structured data built for LLM consumption, handling JavaScript rendering and anti-bot pages.
| Category | Document AI & OCR |
| Type | Web scraping for LLMs |
| License | AGPL-3.0 |
| Runs locally | Self-hosted |
| Built with | TypeScript |
| Skill level | Beginner |
| Best for | developers feeding web content to models |
Other open-source document ai & ocr tools worth comparing:
MarkerConvert PDFs to clean Markdown fast
DoclingIBM-grade document understanding
SuryaModern OCR with layout detection
PaddleOCRIndustrial OCR with tiny models
UnstructuredTurn any file into LLM-ready data
Crawl4AIOpen-source LLM-friendly crawler
olmOCRTurn PDFs into clean training-grade text
MinerUPDF to Markdown with formulas and tables
TesseractThe classic OCR engineFirecrawl is free and open-source (AGPL-3.0 license), so you can use, self-host and modify it at no cost.
Yes. Firecrawl is designed to run on your own machine or server, keeping your data private.
Popular open-source alternatives include Marker, Docling, Surya. See the comparisons above to choose.
Browse the full directory of open-source AI tools, models and projects — updated daily.
Browse all tools →