choose when you want a peer research model for complex document layout benchmarks and can accept an unasserted license.
The official repo for “Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting”, ACL, 2025.
- 9.1k
- Python
- NOASSERTION
document OCR / layout parsing
dots.ocr is a multilingual document layout parsing model that turns document images into structured outputs like markdown and SVG using a single vision-language model.
Same problem, same approach. Swapping one for another is a config change, not a rewrite.
choose when you want a peer research model for complex document layout benchmarks and can accept an unasserted license.
The official repo for “Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting”, ACL, 2025.
choose when you need one-shot parsing of very long documents beyond the seed's input length.
Unlimited OCR Works: Welcome the Era of One-shot Long-horizon Parsing.
choose when you want the GLM-family OCR VLM with Apache-2.0 weights and fast inference.
GLM-OCR: Accurate × Fast × Comprehensive
choose when you need a lightweight LMM for document parsing in memory-constrained environments.
A lightweight LMM-based Document Parsing Model
Solves the same problem with a different architecture or at a different layer. Expect to rewrite the integration.
choose when you need a modular, trainable OCR pipeline with CPU/edge deployment instead of a single VLM.
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
choose when you need full document conversion to multiple formats with metadata and chunking, not just image-to-markdown.
Get your documents ready for gen AI
choose when you want a production ETL system for documents with file-type connectors and vector-database output.
Convert documents to structured data effortlessly. Unstructured is open-source ETL solution for transforming complex documents into clean, structured formats for language models. Visit our website to learn more about our enterprise grade Platform product for production grade workflows, partitioning, enrichments, chunking and embedding.
choose when you prefer a Rust-native parser for various document formats over a Python VLM.
A fast, helpful, and open-source document parser
choose when you need a PyTorch/TensorFlow OCR library you can fine-tune on your own text/layout tasks.
docTR (Document Text Recognition) - a seamless, high-performing & accessible library for OCR-related tasks powered by Deep Learning. Ongoing development and maintenance by t2k.
choose when you want to use hosted vision LLMs (GPT-4o/Claude) for extraction instead of deploying a local model.
OCR & Document Extraction using vision models
choose when you need GPL-licensed unified parsing of documents, audio, and video for GenAI ingestion.
Ingest, parse, and optimize any data format ➡️ from documents to multimedia ➡️ for enhanced compatibility with GenAI frameworks
choose when you need a library with OCR plus layout, reading order, and table recognition in 90+ languages.
OCR, layout analysis, reading order, table recognition in 90+ languages
choose when your inputs are born-digital PDFs and you prefer a Java library for AI-ready extraction without OCR.
PDF Parser for AI-ready data. Automate PDF accessibility. Open-source.
An earlier or more limited way to do the job. Still the right call when you need something small, proven, or CPU-only.
choose when you need an older lighter OCR-free VLM plus a synthetic document generator, with MIT license.
Official Implementation of OCR-free Document Understanding Transformer (Donut) and Synthetic Document Generator (SynthDoG), ECCV 2022
These projects were analysed and named dots.ocr among their alternatives. The relationship is not symmetric — how dots.ocr rates them is a separate judgement, made when dots.ocr is analysed in its own right.
calls dots.ocr “VLM multilingual layout parser; pick for an alternative model with different language strengths.”
Unlimited OCR Works: Welcome the Era of One-shot Long-horizon Parsing.
calls dots.ocr “another LMM for multilingual document layout parsing, a direct drop-in alternative”
A lightweight LMM-based Document Parsing Model
calls dots.ocr “if you need a single vision-language model for multilingual layout parsing”
Get your documents ready for gen AI