AltHub

axa-group/Parsr alternatives

document parsing / extraction

Parsr is a document parsing and extraction toolchain that converts PDFs, images, DOCX, and EML files into enriched structured data (JSON, Markdown, CSV, TXT) with hierarchy and label detection.

6.2kJavaScriptApache-2.0active · last push 5mo agoGitHub

Drop-in peers

Same problem, same approach. Swapping one for another is a config change, not a rewrite.

run-llama/liteparse

actively maintained Apache-2.0 Rust parser, a direct drop-in replacement for Parsr.

A fast, helpful, and open-source document parser

12.2k
Rust
Apache-2.0
active2d ago
Unstructured-IO/unstructured

Python ETL library with active development, outputs clean structured data for LLM pipelines.

Convert documents to structured data effortlessly. Unstructured is open-source ETL solution for transforming complex documents into clean, structured formats for language models. Visit our website to learn more about our enterprise grade Platform product for production grade workflows, partitioning, enrichments, chunking and embedding.

15.3k
HTML
Apache-2.0
active2d ago
PaddlePaddle/PaddleOCR

heavyweight OCR toolkit with 100+ languages, good for image/PDF extraction but more resource-intensive.

Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.

88.1k
Python
Apache-2.0
active1mo ago
opendatalab/PDF-Extract-Kit

research-oriented PDF extraction toolkit under AGPL-3.0, similar scope but copyleft license.

A Comprehensive Toolkit for High-Quality PDF Content Extraction

10.0k
Python
AGPL-3.0
dormant1.6y ago
docling-project/docling

mature MIT-licensed Python library, fast and actively developed for document structure extraction.

Get your documents ready for gen AI

65.4k
Python
MIT
active2d ago
opendataloader-project/opendataloader-pdf

Java-based Apache-2.0 PDF parser designed for AI-ready data, includes accessibility features.

PDF Parser for AI-ready data. Automate PDF accessibility. Open-source.

28.6k
Java
Apache-2.0
active2d ago
xberg-io/xberg

MIT-licensed Rust framework extracting from 101 formats, broader format coverage than Parsr.

A polyglot document intelligence framework with a Rust core. Extract text, metadata, images, and structured data from 101 formats (115 file extensions) plus code intelligence for 371 code languages. 15 language bindings — Rust, Python, Ruby, Java, Go, PHP, Elixir, C#, TypeScript — plus CLI, REST API, and MCP server.

9.2k
Rust
MIT
active1d ago
adithya-s-k/omniparse

GPL-3.0 Python parser for documents/multimedia, similar goal but license may be restrictive.

Ingest, parse, and optimize any data format ➡️ from documents to multimedia ➡️ for enhanced compatibility with GenAI frameworks

7.8k
Python
GPL-3.0
slowing8mo ago

Same job, different approach

Solves the same problem with a different architecture or at a different layer. Expect to rewrite the integration.

landing-ai/ade-cli

LLM-based agentic extraction CLI, differs from Parsr's rule-based pipeline.

The official CLI for Agentic Document Extraction (ADE) by LandingAI — parse documents and extract schema-shaped data from your terminal

2.4k
Python
Apache-2.0
active4d ago
bytedance/Dolphin

research model for document image parsing via deep learning, requires GPU and is not a maintained tool.

The official repo for “Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting”, ACL, 2025.

9.1k
Python
NOASSERTION
active5mo ago

Older generation, or narrower

An earlier or more limited way to do the job. Still the right call when you need something small, proven, or CPU-only.

oomol-lab/pdf-craft

narrower, focuses on scanned-book PDF conversion to formats rather than full document structure extraction.

PDF craft can convert PDF files into various other formats. This project will focus on processing PDF files of scanned books.

6.2k
Python
MIT
active1d ago
jsvine/pdfplumber

Python library for detailed text/table extraction, more limited in document structure understanding.

Plumb a PDF for detailed information about each char, rectangle, line, et cetera — and easily extract text and tables.

10.7k
Python
MIT
active17d ago
clovaai/donut

research transformer for OCR-free document parsing, not production-ready and unmaintained.

Official Implementation of OCR-free Document Understanding Transformer (Donut) and Synthetic Document Generator (SynthDoG), ECCV 2022

6.9k
Python
MIT
dormant2.1y ago
atlanhq/camelot

archived PDF table extraction library, narrow use case and not maintained.

Camelot: PDF Table Extraction for Humans

3.7k
Python
NOASSERTION
archived3.6y ago

Listed as an alternative to

These projects were analysed and named Parsr among their alternatives. The relationship is not symmetric — how Parsr rates them is a separate judgement, made when Parsr is analysed in its own right.

PaddlePaddle/PaddleOCR

calls Parsrif you prefer a Node.js pipeline that outputs enriched JSON for downstream processing.

Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.

88.1kPythonactive
opendatalab/MinerU

calls ParsrJavaScript-based document-to-structured-data converter, for JS ecosystems.

Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.

78.2kPythonactive
docling-project/docling

calls Parsrif you want a self-hosted server with a REST API for document structuring

Get your documents ready for gen AI

65.4kPythonactive
datalab-to/marker

calls ParsrWhen you need a self-hosted REST API service for document restructuring, not a library.

Convert PDF to markdown + JSON quickly with high accuracy

39.0kPythonactive
allenai/olmocr

calls ParsrChoose if you want a JavaScript-based document-to-structured-data pipeline with a REST API.

Toolkit for linearizing PDFs for LLM datasets/training

19.4kPythonactive
firecrawl/anydoc

calls ParsrDeployable document-to-JSON server; choose when you want a service instead of an embedded library.

Convert Word, PowerPoint, Excel, OpenDocument, RTF, EPUB, CSV, and PDF to clean Markdown. Built in Rust, with Node.js and Python bindings.

17.8kRustactive
Unstructured-IO/unstructured

calls Parsrwhen you want a self-hosted Node.js service that converts documents to enriched structured data instead of a Python library.

Convert documents to structured data effortlessly. Unstructured is open-source ETL solution for transforming complex documents into clean, structured formats for language models. Visit our website to learn more about our enterprise grade Platform product for production grade workflows, partitioning, enrichments, chunking and embedding.

15.3kHTMLactive
run-llama/liteparse

calls ParsrJavaScript tool converting documents to structured data; choose it for Node.js/server workflows.

A fast, helpful, and open-source document parser

12.2kRustactive
oomol-lab/pdf-craft

calls ParsrExtracts structured data from PDFs/images using traditional document analysis; choose when you need JSON output rather than Markdown/EPUB.

PDF craft can convert PDF files into various other formats. This project will focus on processing PDF files of scanned books.

6.2kPythonactive
Dicklesworthstone/llm_aided_ocr

calls ParsrAlternative document-to-structured-data tool with OCR support.

Enhances Tesseract OCR output using LLMs (local or API) for error correction, smart chunking, and markdown formatting of scanned PDFs

3.0kPythonactive
NanoNets/docstrange

calls Parsrchoose as a legacy on-prem pipeline for rule-based extraction; it's less active and lacks modern LLM-Markdown output.

Extract and convert data from any document, images, pdfs, word doc, ppt or URL into multiple formats (Markdown, JSON, CSV, HTML) with intelligent structured data extraction and advanced OCR.

1.5kPythonslowing
opendataloader-project/opendataloader-pdf

calls ParsrChoose if you want a standalone processing server with a UI for visualizing extracted data; good for exploratory work, not embedding.

PDF Parser for AI-ready data. Automate PDF accessibility. Open-source.

28.6kJavaactive
getomni-ai/zerox

calls Parsrwhen you want to deploy a server-based document parsing service with a GUI/API.

OCR & Document Extraction using vision models

12.3kTypeScriptslowing
opendatalab/PDF-Extract-Kit

calls Parsrwhen you want a self-hosted document parsing service with a UI/API

A Comprehensive Toolkit for High-Quality PDF Content Extraction

10.0kPythondormant