When you want an IBM-backed, MIT-licensed converter with strong support for complex document structures.
Get your documents ready for gen AI
- 65.4k
- Python
- MIT
document conversion / PDF-to-markdown
Converts PDFs, images, and office documents into markdown, JSON, chunks, and HTML with high accuracy using Datalab's deep learning models.
Same problem, same approach. Swapping one for another is a config change, not a rewrite.
When you want an IBM-backed, MIT-licensed converter with strong support for complex document structures.
Get your documents ready for gen AI
When you need a direct drop-in replacement with active development and easy install.
Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
When you need an ETL pipeline with connectors to many data sources and output formats.
Convert documents to structured data effortlessly. Unstructured is open-source ETL solution for transforming complex documents into clean, structured formats for language models. Visit our website to learn more about our enterprise grade Platform product for production grade workflows, partitioning, enrichments, chunking and embedding.
When you want a lightweight parser focused on lossless LLM ingestion.
File Parser optimised for LLM Ingestion with no loss 🧠 Parse PDFs, Docx, PPTx in a format that is ideal for LLMs.
When you need small models and excellent math formula recognition to Markdown.
An Open-Source Python3 tool with SMALL models for recognizing layouts, tables, math formulas (LaTeX), and text in images, converting them into Markdown format. A free alternative to Mathpix, empowering seamless conversion of visual content into text-based representations. 80+ languages are supported.
When you want a simple MIT-licensed parser with good visual table extraction.
Improved file parsing for LLM’s
When you want a vision-LLM-based converter that can run locally or via API.
A high-quality PDF to Markdown tool based on large language model visual recognition. 一款基于大模型视觉识别的高质量PDF转Markdown工具
When you want an OCR-free, on-prem toolkit with a built-in benchmark suite.
An on-premises, OCR-free unstructured data extraction, markdown conversion and benchmarking toolkit. (https://idp-leaderboard.org/)
When you need multi-format extraction with both OCR and no-OCR paths.
Extract and convert data from any document, images, pdfs, word doc, ppt or URL into multiple formats (Markdown, JSON, CSV, HTML) with intelligent structured data extraction and advanced OCR.
When you want a lightweight multimodal toolkit with cross-page table handling.
OCRFlux is a lightweight yet powerful multimodal toolkit that significantly advances PDF-to-Markdown conversion, excelling in complex layout handling, complicated table parsing and cross-page content merging.
Solves the same problem with a different architecture or at a different layer. Expect to rewrite the integration.
When you are in a Java/JVM environment and need PDF parsing for AI data.
PDF Parser for AI-ready data. Automate PDF accessibility. Open-source.
When you need a battle-tested OCR engine with 100+ languages and full pipeline control.
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
When you want to leverage vision LLMs (GPT-4V, etc.) with a simple TypeScript API.
OCR & Document Extraction using vision models
When you want a Rust-core parser with Python bindings for high throughput.
A fast, helpful, and open-source document parser
When you need a self-hosted REST API service for document restructuring, not a library.
Transforms PDF, Documents and Images into Enriched Structured Data
When you need a polyglot extraction framework covering 101 formats and code extraction.
A polyglot document intelligence framework with a Rust core. Extract text, metadata, images, and structured data from 101 formats (115 file extensions) plus code intelligence for 371 code languages. 15 language bindings — Rust, Python, Ruby, Java, Go, PHP, Elixir, C#, TypeScript — plus CLI, REST API, and MCP server.
When you want a modular framework to build custom document AI pipelines with fine-grained control.
A Repo For Document AI
When you need a RAG-focused toolkit that also handles chunking, cleaning, and metadata enrichment.
Open-source toolkit for reliable RAG pipelines: convert PDFs to Markdown, clean documents, inspect chunks, compare chunking strategies, and enrich metadata for LLM applications.
When you want a Rust-native, multi-format converter with Node.js and Python bindings.
Convert Word, PowerPoint, Excel, OpenDocument, RTF, EPUB, CSV, and PDF to clean Markdown. Built in Rust, with Node.js and Python bindings.
An earlier or more limited way to do the job. Still the right call when you need something small, proven, or CPU-only.
When you only need component extraction (text, formulas, tables) and want to assemble your own pipeline.
A Comprehensive Toolkit for High-Quality PDF Content Extraction
When you need extremely fast text extraction from simple born-digital PDFs.
The fastest PDF library for Python and Rust. Text extraction, image extraction, markdown conversion, PDF creation & editing. 0.8ms mean, 5× faster than industry leaders, 100% pass rate on 3,830 PDFs. MIT/Apache-2.0.
When you want to apply a specific research model for document image parsing.
The official repo for “Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting”, ACL, 2025.
When your task is only PDF table extraction, not full document conversion.
Camelot: PDF Table Extraction for Humans
These projects were analysed and named marker among their alternatives. The relationship is not symmetric — how marker rates them is a separate judgement, made when marker is analysed in its own right.
calls marker “Choose if you want a widely used, Apache-2.0 converter with high accuracy and active development.”
Toolkit for linearizing PDFs for LLM datasets/training
calls marker “drop-in peer for PDF-to-Markdown conversion with high accuracy and Apache-2.0 license”
A Comprehensive Toolkit for High-Quality PDF Content Extraction
calls marker “choose when you need a fast, accurate PDF-to-Markdown converter for published documents; it's lighter than seed but PDF-only.”
Extract and convert data from any document, images, pdfs, word doc, ppt or URL into multiple formats (Markdown, JSON, CSV, HTML) with intelligent structured data extraction and advanced OCR.
calls marker “Python deep-learning PDF-to-markdown; choose it for high accuracy on complex layouts, accepting heavier runtime.”
A fast, helpful, and open-source document parser
calls marker “PDF-only converter to markdown/JSON with high accuracy; choose for PDF processing where marker's accuracy is required.”
Python tool for converting files and office documents to Markdown.
calls marker “if you only need high-accuracy PDF-to-Markdown/JSON conversion”
Get your documents ready for gen AI