another LMM for multilingual document layout parsing, a direct drop-in alternative
Multilingual Document Layout Parsing in a Single Vision-Language Model
- 9.1k
- Python
- MIT
document OCR / parsing
MonkeyOCR is a lightweight LMM-based document parsing model that converts document images into structured text using a structure-recognition-relation triplet paradigm.
Same problem, same approach. Swapping one for another is a config change, not a rewrite.
another LMM for multilingual document layout parsing, a direct drop-in alternative
Multilingual Document Layout Parsing in a Single Vision-Language Model
LMM-based document image parser, similar capability to MonkeyOCR
The official repo for “Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting”, ACL, 2025.
OCR-free document understanding transformer, same LMM-based approach
Official Implementation of OCR-free Document Understanding Transformer (Donut) and Synthetic Document Generator (SynthDoG), ECCV 2022
LMM-based one-shot long-horizon OCR, alternative document parser
Unlimited OCR Works: Welcome the Era of One-shot Long-horizon Parsing.
Solves the same problem with a different architecture or at a different layer. Expect to rewrite the integration.
classic OCR toolkit with broad language support, different architecture from LMM
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
document conversion pipeline for gen AI, not LMM-based but same job
Get your documents ready for gen AI
fast document parser in Rust, different implementation approach
A fast, helpful, and open-source document parser
ETL for transforming documents into structured formats for LLMs
Convert documents to structured data effortlessly. Unstructured is open-source ETL solution for transforming complex documents into clean, structured formats for language models. Visit our website to learn more about our enterprise grade Platform product for production grade workflows, partitioning, enrichments, chunking and embedding.
deep learning OCR library focused on text recognition
docTR (Document Text Recognition) - a seamless, high-performing & accessible library for OCR-related tasks powered by Deep Learning. Ongoing development and maintenance by t2k.
pixel-native web/document parsing for RAG, different indexing approach
The end of web parsing. The beginning of scalable pixel-native search. link: https://pixelrag.ai/
Does the same job, but ships as an app. Useful to a person, not swappable into a codebase.
end-user document management system that performs OCR internally
A community-supported supercharged document management system: scan, index and archive all your documents
These projects were analysed and named MonkeyOCR among their alternatives. The relationship is not symmetric — how MonkeyOCR rates them is a separate judgement, made when MonkeyOCR is analysed in its own right.
calls MonkeyOCR “Lightweight VLM document parser; choose for lower inference resource requirements.”
Unlimited OCR Works: Welcome the Era of One-shot Long-horizon Parsing.
calls MonkeyOCR “choose when you need a lightweight LMM for document parsing in memory-constrained environments.”
Multilingual Document Layout Parsing in a Single Vision-Language Model
calls MonkeyOCR “if you need a lightweight LMM-based document parsing model for low-resource or on-device use”
Get your documents ready for gen AI