actively maintained Apache-2.0 Rust parser, a direct drop-in replacement for Parsr.
A fast, helpful, and open-source document parser
- 12.2k
- Rust
- Apache-2.0
document parsing / extraction
Parsr is a document parsing and extraction toolchain that converts PDFs, images, DOCX, and EML files into enriched structured data (JSON, Markdown, CSV, TXT) with hierarchy and label detection.
Same problem, same approach. Swapping one for another is a config change, not a rewrite.
actively maintained Apache-2.0 Rust parser, a direct drop-in replacement for Parsr.
A fast, helpful, and open-source document parser
Python ETL library with active development, outputs clean structured data for LLM pipelines.
Convert documents to structured data effortlessly. Unstructured is open-source ETL solution for transforming complex documents into clean, structured formats for language models. Visit our website to learn more about our enterprise grade Platform product for production grade workflows, partitioning, enrichments, chunking and embedding.
heavyweight OCR toolkit with 100+ languages, good for image/PDF extraction but more resource-intensive.
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
research-oriented PDF extraction toolkit under AGPL-3.0, similar scope but copyleft license.
A Comprehensive Toolkit for High-Quality PDF Content Extraction
mature MIT-licensed Python library, fast and actively developed for document structure extraction.
Get your documents ready for gen AI
Java-based Apache-2.0 PDF parser designed for AI-ready data, includes accessibility features.
PDF Parser for AI-ready data. Automate PDF accessibility. Open-source.
MIT-licensed Rust framework extracting from 101 formats, broader format coverage than Parsr.
A polyglot document intelligence framework with a Rust core. Extract text, metadata, images, and structured data from 101 formats (115 file extensions) plus code intelligence for 371 code languages. 15 language bindings — Rust, Python, Ruby, Java, Go, PHP, Elixir, C#, TypeScript — plus CLI, REST API, and MCP server.
GPL-3.0 Python parser for documents/multimedia, similar goal but license may be restrictive.
Ingest, parse, and optimize any data format ➡️ from documents to multimedia ➡️ for enhanced compatibility with GenAI frameworks
Solves the same problem with a different architecture or at a different layer. Expect to rewrite the integration.
LLM-based agentic extraction CLI, differs from Parsr's rule-based pipeline.
The official CLI for Agentic Document Extraction (ADE) by LandingAI — parse documents and extract schema-shaped data from your terminal
research model for document image parsing via deep learning, requires GPU and is not a maintained tool.
The official repo for “Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting”, ACL, 2025.
An earlier or more limited way to do the job. Still the right call when you need something small, proven, or CPU-only.
narrower, focuses on scanned-book PDF conversion to formats rather than full document structure extraction.
PDF craft can convert PDF files into various other formats. This project will focus on processing PDF files of scanned books.
Python library for detailed text/table extraction, more limited in document structure understanding.
Plumb a PDF for detailed information about each char, rectangle, line, et cetera — and easily extract text and tables.
research transformer for OCR-free document parsing, not production-ready and unmaintained.
Official Implementation of OCR-free Document Understanding Transformer (Donut) and Synthetic Document Generator (SynthDoG), ECCV 2022
archived PDF table extraction library, narrow use case and not maintained.
Camelot: PDF Table Extraction for Humans
These projects were analysed and named Parsr among their alternatives. The relationship is not symmetric — how Parsr rates them is a separate judgement, made when Parsr is analysed in its own right.
calls Parsr “if you prefer a Node.js pipeline that outputs enriched JSON for downstream processing.”
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
calls Parsr “JavaScript-based document-to-structured-data converter, for JS ecosystems.”
Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
calls Parsr “if you want a self-hosted server with a REST API for document structuring”
Get your documents ready for gen AI
calls Parsr “When you need a self-hosted REST API service for document restructuring, not a library.”
Convert PDF to markdown + JSON quickly with high accuracy
calls Parsr “Choose if you want a JavaScript-based document-to-structured-data pipeline with a REST API.”
Toolkit for linearizing PDFs for LLM datasets/training
calls Parsr “Deployable document-to-JSON server; choose when you want a service instead of an embedded library.”
Convert Word, PowerPoint, Excel, OpenDocument, RTF, EPUB, CSV, and PDF to clean Markdown. Built in Rust, with Node.js and Python bindings.
calls Parsr “when you want a self-hosted Node.js service that converts documents to enriched structured data instead of a Python library.”
Convert documents to structured data effortlessly. Unstructured is open-source ETL solution for transforming complex documents into clean, structured formats for language models. Visit our website to learn more about our enterprise grade Platform product for production grade workflows, partitioning, enrichments, chunking and embedding.
calls Parsr “JavaScript tool converting documents to structured data; choose it for Node.js/server workflows.”
A fast, helpful, and open-source document parser
calls Parsr “Extracts structured data from PDFs/images using traditional document analysis; choose when you need JSON output rather than Markdown/EPUB.”
PDF craft can convert PDF files into various other formats. This project will focus on processing PDF files of scanned books.
calls Parsr “Alternative document-to-structured-data tool with OCR support.”
Enhances Tesseract OCR output using LLMs (local or API) for error correction, smart chunking, and markdown formatting of scanned PDFs
calls Parsr “choose as a legacy on-prem pipeline for rule-based extraction; it's less active and lacks modern LLM-Markdown output.”
Extract and convert data from any document, images, pdfs, word doc, ppt or URL into multiple formats (Markdown, JSON, CSV, HTML) with intelligent structured data extraction and advanced OCR.
calls Parsr “Choose if you want a standalone processing server with a UI for visualizing extracted data; good for exploratory work, not embedding.”
PDF Parser for AI-ready data. Automate PDF accessibility. Open-source.
calls Parsr “when you want to deploy a server-based document parsing service with a GUI/API.”
OCR & Document Extraction using vision models
calls Parsr “when you want a self-hosted document parsing service with a UI/API”
A Comprehensive Toolkit for High-Quality PDF Content Extraction