Direct Python competitor, use when you want an MIT-licensed parser with similar features.
Get your documents ready for gen AI
- 65.4k
- Python
- MIT
document parsing
MinerU is a Python library that transforms PDFs and Office documents into LLM-ready markdown/JSON using layout analysis and OCR.
Same problem, same approach. Swapping one for another is a config change, not a rewrite.
Direct Python competitor, use when you want an MIT-licensed parser with similar features.
Get your documents ready for gen AI
Java-based PDF parser for AI-ready data, use in Java projects.
PDF Parser for AI-ready data. Automate PDF accessibility. Open-source.
Direct Python alternative, use when you want a more established ETL-oriented parser.
Convert documents to structured data effortlessly. Unstructured is open-source ETL solution for transforming complex documents into clean, structured formats for language models. Visit our website to learn more about our enterprise grade Platform product for production grade workflows, partitioning, enrichments, chunking and embedding.
Direct Python alternative, focuses on lossless LLM ingestion.
File Parser optimised for LLM Ingestion with no loss 🧠 Parse PDFs, Docx, PPTx in a format that is ideal for LLMs.
Python library that extracts logical structure and tables, similar to MinerU.
Dedoc is a library (service) for automate documents parsing and bringing to a uniform format. It automatically extracts content, logical structure, tables, and meta information from textual electronic documents. (Parse document; Document content extraction; Logical structure extraction; PDF parser; Scanned document parser; DOCX parser; HTML parser
Solves the same problem with a different architecture or at a different layer. Expect to rewrite the integration.
Rust-based parser, use when you need a fast native binary rather than Python.
A fast, helpful, and open-source document parser
JavaScript-based document-to-structured-data converter, for JS ecosystems.
Transforms PDF, Documents and Images into Enriched Structured Data
An earlier or more limited way to do the job. Still the right call when you need something small, proven, or CPU-only.
A smaller/simpler document parser, suitable for basic extraction tasks.
📄🔍 Parse, extract, and analyze documents with ease 📄🔍
Specializes in Korean documents (HWP/HWPX) to Markdown, use for Korean-specific formats.
모두 파싱해버리겠다 — HWP·HWPX·PDF·Office 문서를 Markdown으로. 양식 자동 채우기와 신구대조를 갖춘 CLI·MCP 서버 | Convert Korean documents (HWP, HWPX, PDF, Office) to Markdown — CLI and MCP server with form filling and diff
Node.js parser for office files only, lacks PDF support.
A robust, strictly-typed Node.js and Browser library for parsing office files into a rich Abstract Syntax Tree (AST) and generating high-fidelity output in multiple formats. Parses: docx · pptx · xlsx · odt · odp · ods · pdf · rtf · csv · md · html. Generates: Markdown · HTML · CSV · RTF · PDF · Plain Text · RAG Chunks
A lightweight LLM-based PDF parser, less complete than MinerU.
A package for parsing PDFs and analyzing their content using LLMs.
Does the same job, but ships as an app. Useful to a person, not swappable into a codebase.
A self-hosted AI document extraction platform, not a drop-in library.
Open-source platform for extracting structured data from documents using AI.
A deployable extraction API service, not importable as a library.
Document (PDF, Word, PPTX ...) extraction and parse API using state of the art modern OCRs + Ollama supported models. Anonymize documents. Remove PII. Convert any document or picture to structured JSON or Markdown
A CLI wrapper for LandingAI's commercial extraction service.
The official CLI for Agentic Document Extraction (ADE) by LandingAI — parse documents and extract schema-shaped data from your terminal
These projects were analysed and named MinerU among their alternatives. The relationship is not symmetric — how MinerU rates them is a separate judgement, made when MinerU is analysed in its own right.
calls MinerU “Python library that converts PDFs and Office docs to LLM-ready markdown/JSON; choose for advanced PDF/OCR-heavy workflows.”
Python tool for converting files and office documents to Markdown.
calls MinerU “when you prefer a parser focused exclusively on document-to-markdown rather than a general OCR toolkit.”
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
calls MinerU “if you want a similarly capable Python parser with a larger community and active development”
Get your documents ready for gen AI
calls MinerU “When you need a direct drop-in replacement with active development and easy install.”
Convert PDF to markdown + JSON quickly with high accuracy
calls MinerU “Choose if you want a popular community-driven Python parser with strong handling of complex PDFs; be aware the license is not clearly OSS.”
PDF Parser for AI-ready data. Automate PDF accessibility. Open-source.
calls MinerU “Choose for a very active, feature-rich option that converts PDFs and Office docs to markdown/JSON.”
Toolkit for linearizing PDFs for LLM datasets/training
calls MinerU “when you want a high-quality PDF/Office-to-Markdown/JSON converter with OCR and table/LaTeX support for LLM workflows.”
Convert documents to structured data effortlessly. Unstructured is open-source ETL solution for transforming complex documents into clean, structured formats for language models. Visit our website to learn more about our enterprise grade Platform product for production grade workflows, partitioning, enrichments, chunking and embedding.
calls MinerU “choose as a drop-in parser for PDF/Office to Markdown/JSON with strong OCR, but you'll handle PII separately.”
Document (PDF, Word, PPTX ...) extraction and parse API using state of the art modern OCRs + Ollama supported models. Anonymize documents. Remove PII. Convert any document or picture to structured JSON or Markdown
calls MinerU “choose when you need high-quality PDF-to-Markdown for scientific papers and mixed content; seed has a different model and cloud option.”
Extract and convert data from any document, images, pdfs, word doc, ppt or URL into multiple formats (Markdown, JSON, CSV, HTML) with intelligent structured data extraction and advanced OCR.
calls MinerU “ML-based parser with OCR for complex or scanned PDFs; choose when you need higher accuracy over raw speed.”
Convert Word, PowerPoint, Excel, OpenDocument, RTF, EPUB, CSV, and PDF to clean Markdown. Built in Rust, with Node.js and Python bindings.
calls MinerU “when you want a comprehensive PDF-to-Markdown pipeline with layout/format support.”
OCR & Document Extraction using vision models
calls MinerU “when you need a mature, feature-rich pipeline for PDF/Office-to-markdown with strong community support.”
The official repo for “Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting”, ACL, 2025.
calls MinerU “Mature pipeline for converting complex PDFs to Markdown/JSON using layout detection and OCR; a strong but internally different alternative.”
PDF craft can convert PDF files into various other formats. This project will focus on processing PDF files of scanned books.
calls MinerU “Use when you need a full document-to-markdown pipeline with layout and table recognition, not just OCR correction.”
Enhances Tesseract OCR output using LLMs (local or API) for error correction, smart chunking, and markdown formatting of scanned PDFs