AltHub

datalab-to/marker alternatives

document conversion / PDF-to-markdown

Converts PDFs, images, and office documents into markdown, JSON, chunks, and HTML with high accuracy using Datalab's deep learning models.

39.0kPythonApache-2.0active · last push 16d agoGitHub

Drop-in peers

Same problem, same approach. Swapping one for another is a config change, not a rewrite.

docling-project/docling

When you want an IBM-backed, MIT-licensed converter with strong support for complex document structures.

Get your documents ready for gen AI

65.4k
Python
MIT
active2d ago
opendatalab/MinerU

When you need a direct drop-in replacement with active development and easy install.

Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.

78.2k
Python
NOASSERTION
active4d ago
Unstructured-IO/unstructured

When you need an ETL pipeline with connectors to many data sources and output formats.

Convert documents to structured data effortlessly. Unstructured is open-source ETL solution for transforming complex documents into clean, structured formats for language models. Visit our website to learn more about our enterprise grade Platform product for production grade workflows, partitioning, enrichments, chunking and embedding.

15.3k
HTML
Apache-2.0
active2d ago
QuivrHQ/MegaParse

When you want a lightweight parser focused on lossless LLM ingestion.

File Parser optimised for LLM Ingestion with no loss 🧠 Parse PDFs, Docx, PPTx in a format that is ideal for LLMs.

7.4k
Python
Apache-2.0
slowing1.5y ago
breezedeus/Pix2Text

When you need small models and excellent math formula recognition to Markdown.

An Open-Source Python3 tool with SMALL models for recognizing layouts, tables, math formulas (LaTeX), and text in images, converting them into Markdown format. A free alternative to Mathpix, empowering seamless conversion of visual content into text-based representations. 80+ languages are supported.

3.2k
Jupyter Notebook
MIT
slowing7mo ago
Filimoa/open-parse

When you want a simple MIT-licensed parser with good visual table extraction.

Improved file parsing for LLM’s

3.2k
Python
MIT
active3mo ago
MarkPDFdown/markpdfdown

When you want a vision-LLM-based converter that can run locally or via API.

A high-quality PDF to Markdown tool based on large language model visual recognition. 一款基于大模型视觉识别的高质量PDF转Markdown工具

2.2k
Python
Apache-2.0
slowing7mo ago
NanoNets/docext

When you want an OCR-free, on-prem toolkit with a built-in benchmark suite.

An on-premises, OCR-free unstructured data extraction, markdown conversion and benchmarking toolkit. (https://idp-leaderboard.org/)

2.1k
Python
Apache-2.0
active5mo ago
NanoNets/docstrange

When you need multi-format extraction with both OCR and no-OCR paths.

Extract and convert data from any document, images, pdfs, word doc, ppt or URL into multiple formats (Markdown, JSON, CSV, HTML) with intelligent structured data extraction and advanced OCR.

1.5k
Python
MIT
slowing10mo ago
chatdoc-com/OCRFlux

When you want a lightweight multimodal toolkit with cross-page table handling.

OCRFlux is a lightweight yet powerful multimodal toolkit that significantly advances PDF-to-Markdown conversion, excelling in complex layout handling, complicated table parsing and cross-page content merging.

2.5k
Python
Apache-2.0
active4mo ago

Same job, different approach

Solves the same problem with a different architecture or at a different layer. Expect to rewrite the integration.

opendataloader-project/opendataloader-pdf

When you are in a Java/JVM environment and need PDF parsing for AI data.

PDF Parser for AI-ready data. Automate PDF accessibility. Open-source.

28.6k
Java
Apache-2.0
active2d ago
PaddlePaddle/PaddleOCR

When you need a battle-tested OCR engine with 100+ languages and full pipeline control.

Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.

88.1k
Python
Apache-2.0
active1mo ago
getomni-ai/zerox

When you want to leverage vision LLMs (GPT-4V, etc.) with a simple TypeScript API.

OCR & Document Extraction using vision models

12.3k
TypeScript
MIT
slowing1.3y ago
run-llama/liteparse

When you want a Rust-core parser with Python bindings for high throughput.

A fast, helpful, and open-source document parser

12.2k
Rust
Apache-2.0
active2d ago
axa-group/Parsr

When you need a self-hosted REST API service for document restructuring, not a library.

Transforms PDF, Documents and Images into Enriched Structured Data

6.2k
JavaScript
Apache-2.0
active5mo ago
xberg-io/xberg

When you need a polyglot extraction framework covering 101 formats and code extraction.

A polyglot document intelligence framework with a Rust core. Extract text, metadata, images, and structured data from 101 formats (115 file extensions) plus code intelligence for 371 code languages. 15 language bindings — Rust, Python, Ruby, Java, Go, PHP, Elixir, C#, TypeScript — plus CLI, REST API, and MCP server.

9.2k
Rust
MIT
active1d ago
deepdoctection/deepdoctection

When you want a modular framework to build custom document AI pipelines with fine-grained control.

A Repo For Document AI

3.2k
Python
Apache-2.0
active7d ago
GiovanniPasq/chunky

When you need a RAG-focused toolkit that also handles chunking, cleaning, and metadata enrichment.

Open-source toolkit for reliable RAG pipelines: convert PDFs to Markdown, clean documents, inspect chunks, compare chunking strategies, and enrich metadata for LLM applications.

167
Python
MIT
active29d ago
firecrawl/anydoc

When you want a Rust-native, multi-format converter with Node.js and Python bindings.

Convert Word, PowerPoint, Excel, OpenDocument, RTF, EPUB, CSV, and PDF to clean Markdown. Built in Rust, with Node.js and Python bindings.

17.8k
Rust
MIT
active2d ago

Older generation, or narrower

An earlier or more limited way to do the job. Still the right call when you need something small, proven, or CPU-only.

opendatalab/PDF-Extract-Kit

When you only need component extraction (text, formulas, tables) and want to assemble your own pipeline.

A Comprehensive Toolkit for High-Quality PDF Content Extraction

10.0k
Python
AGPL-3.0
dormant1.6y ago
yfedoseev/pdf_oxide

When you need extremely fast text extraction from simple born-digital PDFs.

The fastest PDF library for Python and Rust. Text extraction, image extraction, markdown conversion, PDF creation & editing. 0.8ms mean, 5× faster than industry leaders, 100% pass rate on 3,830 PDFs. MIT/Apache-2.0.

966
Rust
Apache-2.0
active1d ago
bytedance/Dolphin

When you want to apply a specific research model for document image parsing.

The official repo for “Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting”, ACL, 2025.

9.1k
Python
NOASSERTION
active5mo ago
atlanhq/camelot

When your task is only PDF table extraction, not full document conversion.

Camelot: PDF Table Extraction for Humans

3.7k
Python
NOASSERTION
archived3.6y ago

Listed as an alternative to

These projects were analysed and named marker among their alternatives. The relationship is not symmetric — how marker rates them is a separate judgement, made when marker is analysed in its own right.

allenai/olmocr

calls markerChoose if you want a widely used, Apache-2.0 converter with high accuracy and active development.

Toolkit for linearizing PDFs for LLM datasets/training

19.4kPythonactive
opendatalab/PDF-Extract-Kit

calls markerdrop-in peer for PDF-to-Markdown conversion with high accuracy and Apache-2.0 license

A Comprehensive Toolkit for High-Quality PDF Content Extraction

10.0kPythondormant
NanoNets/docstrange

calls markerchoose when you need a fast, accurate PDF-to-Markdown converter for published documents; it's lighter than seed but PDF-only.

Extract and convert data from any document, images, pdfs, word doc, ppt or URL into multiple formats (Markdown, JSON, CSV, HTML) with intelligent structured data extraction and advanced OCR.

1.5kPythonslowing
run-llama/liteparse

calls markerPython deep-learning PDF-to-markdown; choose it for high accuracy on complex layouts, accepting heavier runtime.

A fast, helpful, and open-source document parser

12.2kRustactive
microsoft/markitdown

calls markerPDF-only converter to markdown/JSON with high accuracy; choose for PDF processing where marker's accuracy is required.

Python tool for converting files and office documents to Markdown.

175.4kPythonactive
docling-project/docling

calls markerif you only need high-accuracy PDF-to-Markdown/JSON conversion

Get your documents ready for gen AI

65.4kPythonactive