AltHub

allenai/olmocr alternatives

PDF/document to text conversion

Toolkit for converting PDFs and other image-based documents into clean, readable plain text/Markdown for LLM datasets and training.

19.4kPythonApache-2.0active · last push 5mo agoGitHub

Drop-in peers

Same problem, same approach. Swapping one for another is a config change, not a rewrite.

QuivrHQ/MegaParse

Choose for Apache-2.0 parser with broader input support (DOCX/PPTX) and a goal of no-loss parsing.

File Parser optimised for LLM Ingestion with no loss 🧠 Parse PDFs, Docx, PPTx in a format that is ideal for LLMs.

7.4k
Python
Apache-2.0
slowing1.5y ago
datalab-to/marker

Choose if you want a widely used, Apache-2.0 converter with high accuracy and active development.

Convert PDF to markdown + JSON quickly with high accuracy

39.0k
Python
Apache-2.0
active16d ago
opendatalab/MinerU

Choose for a very active, feature-rich option that converts PDFs and Office docs to markdown/JSON.

Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.

78.2k
Python
NOASSERTION
active4d ago
docling-project/docling

Choose if you want an MIT-licensed, highly configurable document conversion library with IBM backing.

Get your documents ready for gen AI

65.4k
Python
MIT
active2d ago
opendatalab/PDF-Extract-Kit

Choose if you want a modular, high-quality extraction toolkit (AGPL-3.0) for custom pipelines.

A Comprehensive Toolkit for High-Quality PDF Content Extraction

10.0k
Python
AGPL-3.0
dormant1.6y ago
opendataloader-project/opendataloader-pdf

Choose if you work in the JVM/Java ecosystem and need a PDF parser tailored for AI-ready data.

PDF Parser for AI-ready data. Automate PDF accessibility. Open-source.

28.6k
Java
Apache-2.0
active2d ago

Same job, different approach

Solves the same problem with a different architecture or at a different layer. Expect to rewrite the integration.

PaddlePaddle/PaddleOCR

Choose if you need multi-language OCR and a traditional pipeline that can be customized for many document types.

Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.

88.1k
Python
Apache-2.0
active1mo ago
axa-group/Parsr

Choose if you want a JavaScript-based document-to-structured-data pipeline with a REST API.

Transforms PDF, Documents and Images into Enriched Structured Data

6.2k
JavaScript
Apache-2.0
active5mo ago

Older generation, or narrower

An earlier or more limited way to do the job. Still the right call when you need something small, proven, or CPU-only.

oomol-lab/pdf-craft

Choose when your corpus is mainly scanned books and you need a simpler, MIT-licensed tool.

PDF craft can convert PDF files into various other formats. This project will focus on processing PDF files of scanned books.

6.2k
Python
MIT
active1d ago

Listed as an alternative to

These projects were analysed and named olmocr among their alternatives. The relationship is not symmetric — how olmocr rates them is a separate judgement, made when olmocr is analysed in its own right.

NanoNets/docstrange

calls olmocrchoose when you need to batch-convert PDFs to plain text for LLM training corpora; seed outputs Markdown/JSON and is not a training-data linearizer.

Extract and convert data from any document, images, pdfs, word doc, ppt or URL into multiple formats (Markdown, JSON, CSV, HTML) with intelligent structured data extraction and advanced OCR.

1.5kPythonslowing