AltHub

studio-dots-ai/dots.ocr alternatives

document OCR / layout parsing

dots.ocr is a multilingual document layout parsing model that turns document images into structured outputs like markdown and SVG using a single vision-language model.

9.1kPythonMITactive · last push 5mo agoGitHub

Drop-in peers

Same problem, same approach. Swapping one for another is a config change, not a rewrite.

bytedance/Dolphin

choose when you want a peer research model for complex document layout benchmarks and can accept an unasserted license.

The official repo for “Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting”, ACL, 2025.

9.1k
Python
NOASSERTION
active5mo ago
baidu/Unlimited-OCR

choose when you need one-shot parsing of very long documents beyond the seed's input length.

Unlimited OCR Works: Welcome the Era of One-shot Long-horizon Parsing.

24.3k
Python
MIT
active25d ago
zai-org/GLM-OCR

choose when you want the GLM-family OCR VLM with Apache-2.0 weights and fast inference.

GLM-OCR: Accurate × Fast × Comprehensive

7.3k
Python
Apache-2.0
active4mo ago
Yuliang-Liu/MonkeyOCR

choose when you need a lightweight LMM for document parsing in memory-constrained environments.

A lightweight LMM-based Document Parsing Model

6.6k
Python
Apache-2.0
active1mo ago

Same job, different approach

Solves the same problem with a different architecture or at a different layer. Expect to rewrite the integration.

PaddlePaddle/PaddleOCR

choose when you need a modular, trainable OCR pipeline with CPU/edge deployment instead of a single VLM.

Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.

88.1k
Python
Apache-2.0
active1mo ago
docling-project/docling

choose when you need full document conversion to multiple formats with metadata and chunking, not just image-to-markdown.

Get your documents ready for gen AI

65.4k
Python
MIT
active2d ago
Unstructured-IO/unstructured

choose when you want a production ETL system for documents with file-type connectors and vector-database output.

Convert documents to structured data effortlessly. Unstructured is open-source ETL solution for transforming complex documents into clean, structured formats for language models. Visit our website to learn more about our enterprise grade Platform product for production grade workflows, partitioning, enrichments, chunking and embedding.

15.3k
HTML
Apache-2.0
active2d ago
run-llama/liteparse

choose when you prefer a Rust-native parser for various document formats over a Python VLM.

A fast, helpful, and open-source document parser

12.2k
Rust
Apache-2.0
active2d ago
mindee/doctr

choose when you need a PyTorch/TensorFlow OCR library you can fine-tune on your own text/layout tasks.

docTR (Document Text Recognition) - a seamless, high-performing & accessible library for OCR-related tasks powered by Deep Learning. Ongoing development and maintenance by t2k.

6.3k
Python
Apache-2.0
active2d ago
getomni-ai/zerox

choose when you want to use hosted vision LLMs (GPT-4o/Claude) for extraction instead of deploying a local model.

OCR & Document Extraction using vision models

12.3k
TypeScript
MIT
slowing1.3y ago
adithya-s-k/omniparse

choose when you need GPL-licensed unified parsing of documents, audio, and video for GenAI ingestion.

Ingest, parse, and optimize any data format ➡️ from documents to multimedia ➡️ for enhanced compatibility with GenAI frameworks

7.8k
Python
GPL-3.0
slowing8mo ago
datalab-to/surya

choose when you need a library with OCR plus layout, reading order, and table recognition in 90+ languages.

OCR, layout analysis, reading order, table recognition in 90+ languages

21.3k
Python
Apache-2.0
active2d ago
opendataloader-project/opendataloader-pdf

choose when your inputs are born-digital PDFs and you prefer a Java library for AI-ready extraction without OCR.

PDF Parser for AI-ready data. Automate PDF accessibility. Open-source.

28.6k
Java
Apache-2.0
active2d ago

Older generation, or narrower

An earlier or more limited way to do the job. Still the right call when you need something small, proven, or CPU-only.

clovaai/donut

choose when you need an older lighter OCR-free VLM plus a synthetic document generator, with MIT license.

Official Implementation of OCR-free Document Understanding Transformer (Donut) and Synthetic Document Generator (SynthDoG), ECCV 2022

6.9k
Python
MIT
dormant2.1y ago

Listed as an alternative to

These projects were analysed and named dots.ocr among their alternatives. The relationship is not symmetric — how dots.ocr rates them is a separate judgement, made when dots.ocr is analysed in its own right.

baidu/Unlimited-OCR

calls dots.ocrVLM multilingual layout parser; pick for an alternative model with different language strengths.

Unlimited OCR Works: Welcome the Era of One-shot Long-horizon Parsing.

24.3kPythonactive
Yuliang-Liu/MonkeyOCR

calls dots.ocranother LMM for multilingual document layout parsing, a direct drop-in alternative

A lightweight LMM-based Document Parsing Model

6.6kPythonactive
docling-project/docling

calls dots.ocrif you need a single vision-language model for multilingual layout parsing

Get your documents ready for gen AI

65.4kPythonactive