AltHub

opendatalab/MinerU alternatives

document parsing

MinerU is a Python library that transforms PDFs and Office documents into LLM-ready markdown/JSON using layout analysis and OCR.

78.2kPythonNOASSERTIONactive · last push 4d agoGitHub

Drop-in peers

Same problem, same approach. Swapping one for another is a config change, not a rewrite.

docling-project/docling

Direct Python competitor, use when you want an MIT-licensed parser with similar features.

Get your documents ready for gen AI

65.4k
Python
MIT
active2d ago
opendataloader-project/opendataloader-pdf

Java-based PDF parser for AI-ready data, use in Java projects.

PDF Parser for AI-ready data. Automate PDF accessibility. Open-source.

28.6k
Java
Apache-2.0
active2d ago
Unstructured-IO/unstructured

Direct Python alternative, use when you want a more established ETL-oriented parser.

Convert documents to structured data effortlessly. Unstructured is open-source ETL solution for transforming complex documents into clean, structured formats for language models. Visit our website to learn more about our enterprise grade Platform product for production grade workflows, partitioning, enrichments, chunking and embedding.

15.3k
HTML
Apache-2.0
active2d ago
QuivrHQ/MegaParse

Direct Python alternative, focuses on lossless LLM ingestion.

File Parser optimised for LLM Ingestion with no loss 🧠 Parse PDFs, Docx, PPTx in a format that is ideal for LLMs.

7.4k
Python
Apache-2.0
slowing1.5y ago
ispras/dedoc

Python library that extracts logical structure and tables, similar to MinerU.

Dedoc is a library (service) for automate documents parsing and bringing to a uniform format. It automatically extracts content, logical structure, tables, and meta information from textual electronic documents. (Parse document; Document content extraction; Logical structure extraction; PDF parser; Scanned document parser; DOCX parser; HTML parser

720
Python
Apache-2.0
active12d ago

Same job, different approach

Solves the same problem with a different architecture or at a different layer. Expect to rewrite the integration.

run-llama/liteparse

Rust-based parser, use when you need a fast native binary rather than Python.

A fast, helpful, and open-source document parser

12.2k
Rust
Apache-2.0
active2d ago
axa-group/Parsr

JavaScript-based document-to-structured-data converter, for JS ecosystems.

Transforms PDF, Documents and Images into Enriched Structured Data

6.2k
JavaScript
Apache-2.0
active5mo ago

Older generation, or narrower

An earlier or more limited way to do the job. Still the right call when you need something small, proven, or CPU-only.

AdemBoukhris457/Doctra

A smaller/simpler document parser, suitable for basic extraction tasks.

📄🔍 Parse, extract, and analyze documents with ease 📄🔍

214
Jupyter Notebook
Apache-2.0
slowing9mo ago
chrisryugj/kordoc

Specializes in Korean documents (HWP/HWPX) to Markdown, use for Korean-specific formats.

모두 파싱해버리겠다 — HWP·HWPX·PDF·Office 문서를 Markdown으로. 양식 자동 채우기와 신구대조를 갖춘 CLI·MCP 서버 | Convert Korean documents (HWP, HWPX, PDF, Office) to Markdown — CLI and MCP server with form filling and diff

1.8k
TypeScript
MIT
active1d ago
harshankur/officeParser

Node.js parser for office files only, lacks PDF support.

A robust, strictly-typed Node.js and Browser library for parsing office files into a rich Abstract Syntax Tree (AST) and generating high-fidelity output in multiple formats. Parses: docx · pptx · xlsx · odt · odp · ods · pdf · rtf · csv · md · html. Generates: Markdown · HTML · CSV · RTF · PDF · Plain Text · RAG Chunks

528
Rich Text Format
MIT
active5d ago
lazyFrogLOL/llmdocparser

A lightweight LLM-based PDF parser, less complete than MinerU.

A package for parsing PDFs and analyzing their content using LLMs.

267
Python
MIT
dormant2.0y ago

End-user tools

Does the same job, but ships as an app. Useful to a person, not swappable into a codebase.

DocumindHQ/documind

A self-hosted AI document extraction platform, not a drop-in library.

Open-source platform for extracting structured data from documents using AI.

1.5k
JavaScript
NOASSERTION
slowing1.3y ago
CatchTheTornado/text-extract-api

A deployable extraction API service, not importable as a library.

Document (PDF, Word, PPTX ...) extraction and parse API using state of the art modern OCRs + Ollama supported models. Anonymize documents. Remove PII. Convert any document or picture to structured JSON or Markdown

3.2k
Python
MIT
slowing9mo ago
landing-ai/ade-cli

A CLI wrapper for LandingAI's commercial extraction service.

The official CLI for Agentic Document Extraction (ADE) by LandingAI — parse documents and extract schema-shaped data from your terminal

2.4k
Python
Apache-2.0
active4d ago

Listed as an alternative to

These projects were analysed and named MinerU among their alternatives. The relationship is not symmetric — how MinerU rates them is a separate judgement, made when MinerU is analysed in its own right.

microsoft/markitdown

calls MinerUPython library that converts PDFs and Office docs to LLM-ready markdown/JSON; choose for advanced PDF/OCR-heavy workflows.

Python tool for converting files and office documents to Markdown.

175.4kPythonactive
PaddlePaddle/PaddleOCR

calls MinerUwhen you prefer a parser focused exclusively on document-to-markdown rather than a general OCR toolkit.

Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.

88.1kPythonactive
docling-project/docling

calls MinerUif you want a similarly capable Python parser with a larger community and active development

Get your documents ready for gen AI

65.4kPythonactive
datalab-to/marker

calls MinerUWhen you need a direct drop-in replacement with active development and easy install.

Convert PDF to markdown + JSON quickly with high accuracy

39.0kPythonactive
opendataloader-project/opendataloader-pdf

calls MinerUChoose if you want a popular community-driven Python parser with strong handling of complex PDFs; be aware the license is not clearly OSS.

PDF Parser for AI-ready data. Automate PDF accessibility. Open-source.

28.6kJavaactive
allenai/olmocr

calls MinerUChoose for a very active, feature-rich option that converts PDFs and Office docs to markdown/JSON.

Toolkit for linearizing PDFs for LLM datasets/training

19.4kPythonactive
Unstructured-IO/unstructured

calls MinerUwhen you want a high-quality PDF/Office-to-Markdown/JSON converter with OCR and table/LaTeX support for LLM workflows.

Convert documents to structured data effortlessly. Unstructured is open-source ETL solution for transforming complex documents into clean, structured formats for language models. Visit our website to learn more about our enterprise grade Platform product for production grade workflows, partitioning, enrichments, chunking and embedding.

15.3kHTMLactive
CatchTheTornado/text-extract-api

calls MinerUchoose as a drop-in parser for PDF/Office to Markdown/JSON with strong OCR, but you'll handle PII separately.

Document (PDF, Word, PPTX ...) extraction and parse API using state of the art modern OCRs + Ollama supported models. Anonymize documents. Remove PII. Convert any document or picture to structured JSON or Markdown

3.2kPythonslowing
NanoNets/docstrange

calls MinerUchoose when you need high-quality PDF-to-Markdown for scientific papers and mixed content; seed has a different model and cloud option.

Extract and convert data from any document, images, pdfs, word doc, ppt or URL into multiple formats (Markdown, JSON, CSV, HTML) with intelligent structured data extraction and advanced OCR.

1.5kPythonslowing
firecrawl/anydoc

calls MinerUML-based parser with OCR for complex or scanned PDFs; choose when you need higher accuracy over raw speed.

Convert Word, PowerPoint, Excel, OpenDocument, RTF, EPUB, CSV, and PDF to clean Markdown. Built in Rust, with Node.js and Python bindings.

17.8kRustactive
getomni-ai/zerox

calls MinerUwhen you want a comprehensive PDF-to-Markdown pipeline with layout/format support.

OCR & Document Extraction using vision models

12.3kTypeScriptslowing
bytedance/Dolphin

calls MinerUwhen you need a mature, feature-rich pipeline for PDF/Office-to-markdown with strong community support.

The official repo for “Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting”, ACL, 2025.

9.1kPythonactive
oomol-lab/pdf-craft

calls MinerUMature pipeline for converting complex PDFs to Markdown/JSON using layout detection and OCR; a strong but internally different alternative.

PDF craft can convert PDF files into various other formats. This project will focus on processing PDF files of scanned books.

6.2kPythonactive
Dicklesworthstone/llm_aided_ocr

calls MinerUUse when you need a full document-to-markdown pipeline with layout and table recognition, not just OCR correction.

Enhances Tesseract OCR output using LLMs (local or API) for error correction, smart chunking, and markdown formatting of scanned PDFs

3.0kPythonactive