AltHub

Unstructured-IO/unstructured alternatives

document parsing / OCR / structured data extraction for LLMs

Unstructured is an open-source ETL tool that converts complex documents (PDF, DOCX, images, etc.) into clean structured formats for LLMs via OCR, layout analysis, and partitioning.

15.3kHTMLApache-2.0active · last push 2d agoGitHub

Drop-in peers

Same problem, same approach. Swapping one for another is a config change, not a rewrite.

NanoNets/docstrange

when you need a library that directly OCRs and converts PDFs/images/Office files to Markdown/JSON/CSV/HTML with MIT license.

Extract and convert data from any document, images, pdfs, word doc, ppt or URL into multiple formats (Markdown, JSON, CSV, HTML) with intelligent structured data extraction and advanced OCR.

1.5k
Python
MIT
slowing10mo ago
docling-project/docling

when you need a mature, MIT-licensed document conversion library with strong PDF/Office support and active maintenance.

Get your documents ready for gen AI

65.4k
Python
MIT
active2d ago
opendatalab/MinerU

when you want a high-quality PDF/Office-to-Markdown/JSON converter with OCR and table/LaTeX support for LLM workflows.

Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.

78.2k
Python
NOASSERTION
active4d ago

Same job, different approach

Solves the same problem with a different architecture or at a different layer. Expect to rewrite the integration.

enoch3712/ExtractThinker

when you want ORM-style, LLM-driven schema-based extraction rather than generic document partitioning.

ExtractThinker is a Document Intelligence library for LLMs, offering ORM-style interaction for flexible and powerful document workflows.

1.6k
Python
Apache-2.0
slowing12mo ago
yobix-ai/extractous

when you need Rust-core performance and bindings for multiple languages instead of a Python-native parser.

Fast and efficient unstructured data extraction. Written in Rust with bindings for many languages.

1.8k
Rust
Apache-2.0
dormant1.7y ago
axa-group/Parsr

when you want a self-hosted Node.js service that converts documents to enriched structured data instead of a Python library.

Transforms PDF, Documents and Images into Enriched Structured Data

6.2k
JavaScript
Apache-2.0
active5mo ago
run-llama/liteparse

when you need a fast Rust-based document parser with a minimal API from the LlamaIndex ecosystem.

A fast, helpful, and open-source document parser

12.2k
Rust
Apache-2.0
active2d ago
PaddlePaddle/PaddleOCR

when you want an OCR-centric toolkit with huge language coverage and layout/table recognition rather than a full ETL pipeline.

Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.

88.1k
Python
Apache-2.0
active1mo ago
getomni-ai/zerox

when you want vision-model-based OCR and extraction for complex documents instead of traditional segmentation/parsing.

OCR & Document Extraction using vision models

12.3k
TypeScript
MIT
slowing1.3y ago
deepdoctection/deepdoctection

when you need a customizable document AI framework for building your own extraction pipelines with DL models.

A Repo For Document AI

3.2k
Python
Apache-2.0
active7d ago
marieai/marie-ai

when you need an orchestration framework that integrates LLM/VLLM pipelines for complex document extraction.

Complex data extraction and orchestration framework designed for processing unstructured documents. It integrates AI-powered document pipelines (GenAI, LLM, VLLM) into your applications, supporting various tasks such as document cleanup, optical character recognition (OCR), classification, splitting, named entity recognition, and form processing

95
Python
MIT
active2d ago

Older generation, or narrower

An earlier or more limited way to do the job. Still the right call when you need something small, proven, or CPU-only.

stanford-oval/Churro

when you need specialized OCR for historical/handwritten documents rather than general-purpose parsing.

CHURRO is an OCR toolkit for historical document transcription, built to make handwritten and printed sources readable at high accuracy and lower cost.

68
Python
Apache-2.0
active4mo ago
papercast-dev/papercast

when you need a pipeline focused on academic/technical papers with arXiv and GROBID integration.

A Python pipeline tool and plugin ecosystem for processing technical documents. Process papers from arXiv, SemanticScholar, PDF, with GROBID, LangChain, listen as podcast. Customize your own pipelines.

52
Python
MIT
slowing1.4y ago

End-user tools

Does the same job, but ships as an app. Useful to a person, not swappable into a codebase.

ComPDF-derek/docslight

when you want a ready-made web/UI tool for document parsing rather than embedding a library.

Part of the KDAN ecosystem, DocSlight offers document parsing, OCR, and data extraction that turn PDFs, scans, images, and Office files into structured outputs for RAG pipelines, AI agents, and enterprise document automation.

51
Vue
LGPL-3.0
active23d ago
ComPDFKit/docslight

when you want a hosted/GUI document extraction tool from the ComPDF ecosystem instead of a code library.

With DocSlight, precisely parse and extract data from any document, including PDFs, scans, images, and Office files. It is an open-source AI project from ComPDF (KDAN ecosystem).

123
Vue
LGPL-3.0
active23d ago
Nebutra/MinerU-Skill

when you want a zero-dependency CLI or agent skill that wraps a parser for easy use, not a library to integrate.

AI-Native document parser: PDF, Office & images → clean Markdown with LaTeX, tables & OCR. Zero-dependency CLI & skill for Claude Code, Cursor & AI agents.

102
Python
MIT
active29d ago

Listed as an alternative to

These projects were analysed and named unstructured among their alternatives. The relationship is not symmetric — how unstructured rates them is a separate judgement, made when unstructured is analysed in its own right.

PaddlePaddle/PaddleOCR

calls unstructuredwhen you want a full ETL platform with connectors, API, and pre-built chunking.

Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.

88.1kPythonactive
opendatalab/MinerU

calls unstructuredDirect Python alternative, use when you want a more established ETL-oriented parser.

Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.

78.2kPythonactive
docling-project/docling

calls unstructuredif you need an ETL-oriented document processing pipeline with data connectors

Get your documents ready for gen AI

65.4kPythonactive
datalab-to/marker

calls unstructuredWhen you need an ETL pipeline with connectors to many data sources and output formats.

Convert PDF to markdown + JSON quickly with high accuracy

39.0kPythonactive
opendataloader-project/opendataloader-pdf

calls unstructuredChoose if you need an ETL pipeline with connectors for many document types; it is heavier and may require separate processing for accessibility.

PDF Parser for AI-ready data. Automate PDF accessibility. Open-source.

28.6kJavaactive
firecrawl/anydoc

calls unstructuredPython ETL for document structuring; choose when you need data pipeline integrations.

Convert Word, PowerPoint, Excel, OpenDocument, RTF, EPUB, CSV, and PDF to clean Markdown. Built in Rust, with Node.js and Python bindings.

17.8kRustactive
axa-group/Parsr

calls unstructuredPython ETL library with active development, outputs clean structured data for LLM pipelines.

Transforms PDF, Documents and Images into Enriched Structured Data

6.2kJavaScriptactive
NanoNets/docstrange

calls unstructuredchoose when you need a production-grade ETL with API connectors, partitioning, and LangChain/vector-store integration; seed is simpler but less ecosystem-connected.

Extract and convert data from any document, images, pdfs, word doc, ppt or URL into multiple formats (Markdown, JSON, CSV, HTML) with intelligent structured data extraction and advanced OCR.

1.5kPythonslowing
baidu/Unlimited-OCR

calls unstructuredDocument ETL pipeline; choose if you want connectors and chunking for LLM ingestion.

Unlimited OCR Works: Welcome the Era of One-shot Long-horizon Parsing.

24.3kPythonactive
getomni-ai/zerox

calls unstructuredwhen you need a full ETL pipeline for documents with pluggable backends (not vision-only).

OCR & Document Extraction using vision models

12.3kTypeScriptslowing
studio-dots-ai/dots.ocr

calls unstructuredchoose when you want a production ETL system for documents with file-type connectors and vector-database output.

Multilingual Document Layout Parsing in a Single Vision-Language Model

9.1kPythonactive
clovaai/donut

calls unstructuredchoose when you want an ETL pipeline to turn documents into structured data for LLMs, using a broader toolset rather than an OCR-free model.

Official Implementation of OCR-free Document Understanding Transformer (Donut) and Synthetic Document Generator (SynthDoG), ECCV 2022

6.9kPythondormant
Yuliang-Liu/MonkeyOCR

calls unstructuredETL for transforming documents into structured formats for LLMs

A lightweight LMM-based Document Parsing Model

6.6kPythonactive
oomol-lab/pdf-craft

calls unstructuredETL for document structuring with multiple OCR backends; choose when you need to handle many document types beyond just PDFs.

PDF craft can convert PDF files into various other formats. This project will focus on processing PDF files of scanned books.

6.2kPythonactive
Dicklesworthstone/llm_aided_ocr

calls unstructuredUse for ETL-style document processing into structured data for LLMs.

Enhances Tesseract OCR output using LLMs (local or API) for error correction, smart chunking, and markdown formatting of scanned PDFs

3.0kPythonactive