AltHub

PaddlePaddle/PaddleOCR alternatives

document OCR / parsing

PaddleOCR is a comprehensive OCR and document parsing toolkit that converts PDFs and images into structured, LLM-ready data (JSON/Markdown), supporting 100+ languages.

88.1kPythonApache-2.0active · last push 1mo agoGitHub

Drop-in peers

Same problem, same approach. Swapping one for another is a config change, not a rewrite.

opendatalab/MinerU

when you prefer a parser focused exclusively on document-to-markdown rather than a general OCR toolkit.

Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.

78.2k
Python
NOASSERTION
active4d ago
NanoNets/docstrange

when you need a lightweight Python API that also emits CSV/HTML for downstream apps.

Extract and convert data from any document, images, pdfs, word doc, ppt or URL into multiple formats (Markdown, JSON, CSV, HTML) with intelligent structured data extraction and advanced OCR.

1.5k
Python
MIT
slowing10mo ago
run-llama/liteparse

if you want a fast Rust-based parser with a small footprint and Apache-2.0.

A fast, helpful, and open-source document parser

12.2k
Rust
Apache-2.0
active2d ago
opendataloader-project/opendataloader-pdf

if you need a Java/JVM parser integrated into JVM data pipelines.

PDF Parser for AI-ready data. Automate PDF accessibility. Open-source.

28.6k
Java
Apache-2.0
active2d ago
Topdu/OpenOCR

if you want an open-source toolkit with unified training/evaluation and reproducible OCR baselines.

OpenOCR: An Open-Source Toolkit for General-OCR Research and Applications, integrates a unified training and evaluation benchmark, commercial-grade OCR and Document Parsing systems, and faithful reproductions of the core implementations from a wide range of academic papers.

1.4k
Python
Apache-2.0
active19d ago
axa-group/Parsr

if you prefer a Node.js pipeline that outputs enriched JSON for downstream processing.

Transforms PDF, Documents and Images into Enriched Structured Data

6.2k
JavaScript
Apache-2.0
active5mo ago
Unstructured-IO/unstructured

when you want a full ETL platform with connectors, API, and pre-built chunking.

Convert documents to structured data effortlessly. Unstructured is open-source ETL solution for transforming complex documents into clean, structured formats for language models. Visit our website to learn more about our enterprise grade Platform product for production grade workflows, partitioning, enrichments, chunking and embedding.

15.3k
HTML
Apache-2.0
active2d ago
opendatalab/PDF-Extract-Kit

when you want an AGPL-licensed modular PDF extraction toolkit with high-quality layout outputs.

A Comprehensive Toolkit for High-Quality PDF Content Extraction

10.0k
Python
AGPL-3.0
dormant1.6y ago

Same job, different approach

Solves the same problem with a different architecture or at a different layer. Expect to rewrite the integration.

DocSlicer/DocSlicer

when you need deterministic rule-based parsing/chunking for RAG with an MCP server.

Fast, deterministic document parser and chunker for RAG and agent pipelines. Ships an MCP server for agents navigating long documents with precision.

33
Python
AGPL-3.0
active4d ago
getomni-ai/zerox

when you want to use vision-language models for OCR and can call external model APIs.

OCR & Document Extraction using vision models

12.3k
TypeScript
MIT
slowing1.3y ago
ENDEVSOLS/LongParser

when you need a self-hosted document parsing server with HITL review and multi-format ingestion.

Privacy-first document intelligence engine — parse PDFs, DOCX, PPTX, XLSX & CSV into AI-ready chunks for RAG pipelines. Includes HITL review, 3-layer memory chat, and a production FastAPI server.

30
Python
NOASSERTION
active4mo ago

Older generation, or narrower

An earlier or more limited way to do the job. Still the right call when you need something small, proven, or CPU-only.

opendatalab/MinerU-Diffusion

a research framework for diffusion-based OCR decoding, not a full document parser.

[ECCV 2026] A diffusion-based framework for document OCR that replaces autoregressive decoding with block-level parallel diffusion decoding.

614
Python
MIT
active2mo ago
CycloneBoy/pdf_table

when you only need table extraction and want a dedicated lightweight model.

A Unified Toolkit for Deep Learning-Based Table Extraction

61
Python
none
dormant1.8y ago
jiangnanboy/Image_KIE_LLM

when your task is key-info extraction from cards/bills using LLM prompts, not full-page OCR.

利用llm大语言模型提取卡证票据关键信息。Key Information Extraction from Image with LLM(large language model).Basically, it can extract key information from all bills and documents.

15
Python
MIT
dormant2.1y ago
stanford-oval/Churro

when you are transcribing historical handwritten/printed documents with specialized OCR models.

CHURRO is an OCR toolkit for historical document transcription, built to make handwritten and printed sources readable at high accuracy and lower cost.

68
Python
Apache-2.0
active4mo ago

End-user tools

Does the same job, but ships as an app. Useful to a person, not swappable into a codebase.

ComPDF-derek/docslight

a Vue-based document parsing app/service from KDAN, not a code-embeddable library.

Part of the KDAN ecosystem, DocSlight offers document parsing, OCR, and data extraction that turn PDFs, scans, images, and Office files into structured outputs for RAG pipelines, AI agents, and enterprise document automation.

51
Vue
LGPL-3.0
active23d ago
ComPDFKit/docslight

a web-based document parsing demo/app, not a library.

With DocSlight, precisely parse and extract data from any document, including PDFs, scans, images, and Office files. It is an open-source AI project from ComPDF (KDAN ecosystem).

123
Vue
LGPL-3.0
active23d ago
hiroi-sora/Umi-OCR

a desktop OCR GUI app for end users, not a library for developers.

OCR software, free and offline. 开源、免费的离线OCR软件。支持截屏/批量导入图片,PDF文档识别,排除水印/页眉页脚,扫描/生成二维码。内置多国语言库。

46.7k
Python
MIT
slowing9mo ago

Listed as an alternative to

These projects were analysed and named PaddleOCR among their alternatives. The relationship is not symmetric — how PaddleOCR rates them is a separate judgement, made when PaddleOCR is analysed in its own right.

tesseract-ocr/tesseract

calls PaddleOCRactive Python/PP-OCR toolkit with 100+ languages and PDF-to-structured-data; choose for broad language support plus OCR and document parsing.

Tesseract Open Source OCR Engine (main repository)

76.1kC++active
JaidedAI/EasyOCR

calls PaddleOCRChoose when you need a more production-ready OCR toolkit with built-in table/layout recognition and a larger model zoo, but you are comfortable with PaddlePaddle instead of PyTorch.

Ready-to-use OCR with 80+ supported languages and all popular writing scripts including Latin, Chinese, Arabic, Devanagari, Cyrillic and etc.

29.9kPythonslowing
datalab-to/surya

calls PaddleOCRdrop-in peer with mature, widely-adopted OCR/layout/table toolkit and extensive language support.

OCR, layout analysis, reading order, table recognition in 90+ languages

21.3kPythonactive
datalab-to/chandra

calls PaddleOCRchoose for up to 100+ languages and a mature Apache-2.0 OCR ecosystem.

OCR model that handles complex tables, forms, handwriting with full layout.

12.1kPythonactive
zai-org/GLM-OCR

calls PaddleOCRChoose for a production-ready OCR toolkit with 100+ languages and extensive deployment options.

GLM-OCR: Accurate × Fast × Comprehensive

7.3kPythonactive
axa-group/Parsr

calls PaddleOCRheavyweight OCR toolkit with 100+ languages, good for image/PDF extraction but more resource-intensive.

Transforms PDF, Documents and Images into Enriched Structured Data

6.2kJavaScriptactive
Topdu/OpenOCR

calls PaddleOCRFull-featured OCR toolkit with broad language support and pretrained models, a mature drop-in peer for OpenOCR.

OpenOCR: An Open-Source Toolkit for General-OCR Research and Applications, integrates a unified training and evaluation benchmark, commercial-grade OCR and Document Parsing systems, and faithful reproductions of the core implementations from a wide range of academic papers.

1.4kPythonactive
datalab-to/marker

calls PaddleOCRWhen you need a battle-tested OCR engine with 100+ languages and full pipeline control.

Convert PDF to markdown + JSON quickly with high accuracy

39.0kPythonactive
opendataloader-project/opendataloader-pdf

calls PaddleOCRChoose if you need a battle-tested OCR library with hundreds of languages and fine control; it provides text boxes but not full Markdown/JSON conversion.

PDF Parser for AI-ready data. Automate PDF accessibility. Open-source.

28.6kJavaactive
baidu/Unlimited-OCR

calls PaddleOCRDetection-recognition pipeline toolkit; choose if you need 100+ language support and deployable tools.

Unlimited OCR Works: Welcome the Era of One-shot Long-horizon Parsing.

24.3kPythonactive
allenai/olmocr

calls PaddleOCRChoose if you need multi-language OCR and a traditional pipeline that can be customized for many document types.

Toolkit for linearizing PDFs for LLM datasets/training

19.4kPythonactive
firecrawl/anydoc

calls PaddleOCROCR engine for images and scanned docs; choose when input is raster, not digital.

Convert Word, PowerPoint, Excel, OpenDocument, RTF, EPUB, CSV, and PDF to clean Markdown. Built in Rust, with Node.js and Python bindings.

17.8kRustactive
Unstructured-IO/unstructured

calls PaddleOCRwhen you want an OCR-centric toolkit with huge language coverage and layout/table recognition rather than a full ETL pipeline.

Convert documents to structured data effortlessly. Unstructured is open-source ETL solution for transforming complex documents into clean, structured formats for language models. Visit our website to learn more about our enterprise grade Platform product for production grade workflows, partitioning, enrichments, chunking and embedding.

15.3kHTMLactive
getomni-ai/zerox

calls PaddleOCRwhen you need a powerful multi-language OCR toolkit with traditional layout analysis.

OCR & Document Extraction using vision models

12.3kTypeScriptslowing
studio-dots-ai/dots.ocr

calls PaddleOCRchoose when you need a modular, trainable OCR pipeline with CPU/edge deployment instead of a single VLM.

Multilingual Document Layout Parsing in a Single Vision-Language Model

9.1kPythonactive
Yuliang-Liu/MonkeyOCR

calls PaddleOCRclassic OCR toolkit with broad language support, different architecture from LMM

A lightweight LMM-based Document Parsing Model

6.6kPythonactive
Dicklesworthstone/llm_aided_ocr

calls PaddleOCRChoose for an all-in-one OCR toolkit with more languages and layout awareness than Tesseract alone.

Enhances Tesseract OCR output using LLMs (local or API) for error correction, smart chunking, and markdown formatting of scanned PDFs

3.0kPythonactive
NanoNets/docstrange

calls PaddleOCRchoose when you need lightweight OCR with 100+ languages and CPU/edge deployment; seed uses a heavier 7B model.

Extract and convert data from any document, images, pdfs, word doc, ppt or URL into multiple formats (Markdown, JSON, CSV, HTML) with intelligent structured data extraction and advanced OCR.

1.5kPythonslowing
docling-project/docling

calls PaddleOCRif you primarily need OCR for 100+ languages on images/PDFs, not full-format parsing

Get your documents ready for gen AI

65.4kPythonactive