choose when you need a mature, format-agnostic Python library with deeper pipeline control and custom model support; seed offers a simpler all-in-one 7B approach.
Get your documents ready for gen AI
- 65.4k
- Python
- MIT
document parsing / image-to-markdown / structured extraction
DocStrange converts PDFs, images, Office docs, and URLs into Markdown/JSON/CSV/HTML with OCR and schema-based structured extraction, using a local 7B model or cloud API.
Same problem, same approach. Swapping one for another is a config change, not a rewrite.
choose when you need a mature, format-agnostic Python library with deeper pipeline control and custom model support; seed offers a simpler all-in-one 7B approach.
Get your documents ready for gen AI
choose when you need a fast, embeddable Rust document parser for low-latency services; seed is Python and includes a UI/API.
A fast, helpful, and open-source document parser
choose when you need a production-grade ETL with API connectors, partitioning, and LangChain/vector-store integration; seed is simpler but less ecosystem-connected.
Convert documents to structured data effortlessly. Unstructured is open-source ETL solution for transforming complex documents into clean, structured formats for language models. Visit our website to learn more about our enterprise grade Platform product for production grade workflows, partitioning, enrichments, chunking and embedding.
choose when you need high-quality PDF-to-Markdown for scientific papers and mixed content; seed has a different model and cloud option.
Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
choose when you need a fast, accurate PDF-to-Markdown converter for published documents; it's lighter than seed but PDF-only.
Convert PDF to markdown + JSON quickly with high accuracy
choose when you're in the TypeScript/JavaScript ecosystem and want to use OpenAI-compatible vision models for OCR; seed is Python and has a local model.
OCR & Document Extraction using vision models
choose when you want a Python library using GPT-4o-class vision models for accurate PDF-to-Markdown, accepting API cost and no local mode.
Improved file parsing for LLM’s
choose when you need small CPU-friendly models and math formula support to Markdown; seed's 7B model is heavier.
An Open-Source Python3 tool with SMALL models for recognizing layouts, tables, math formulas (LaTeX), and text in images, converting them into Markdown format. A free alternative to Mathpix, empowering seamless conversion of visual content into text-based representations. 80+ languages are supported.
Solves the same problem with a different architecture or at a different layer. Expect to rewrite the integration.
choose when you need a JVM-native PDF parser to integrate into Java/Scala applications; seed is Python-centric.
PDF Parser for AI-ready data. Automate PDF accessibility. Open-source.
choose when you need lightweight OCR with 100+ languages and CPU/edge deployment; seed uses a heavier 7B model.
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
choose when you want to assemble your own PDF extraction modules (layout, formula, table) under AGPL; seed is an all-in-one tool.
A Comprehensive Toolkit for High-Quality PDF Content Extraction
choose when you need to build a custom document AI pipeline with your own models; not a ready-to-use converter.
A Repo For Document AI
An earlier or more limited way to do the job. Still the right call when you need something small, proven, or CPU-only.
choose when your workload is scanned book PDFs and you need specialized book layout handling; seed covers more formats but with less focus.
PDF craft can convert PDF files into various other formats. This project will focus on processing PDF files of scanned books.
choose as a legacy on-prem pipeline for rule-based extraction; it's less active and lacks modern LLM-Markdown output.
Transforms PDF, Documents and Images into Enriched Structured Data
choose when you need to batch-convert PDFs to plain text for LLM training corpora; seed outputs Markdown/JSON and is not a training-data linearizer.
Toolkit for linearizing PDFs for LLM datasets/training
choose when you need fast Markdown from digital (non-scanned) Office/PDF files with Node/Python bindings; it lacks OCR.
Convert Word, PowerPoint, Excel, OpenDocument, RTF, EPUB, CSV, and PDF to clean Markdown. Built in Rust, with Node.js and Python bindings.
choose when you specifically need .docx to HTML in JavaScript with a tiny dependency; seed's converter spans more formats.
Convert Word documents (.docx files) to HTML
Does the same job, but ships as an app. Useful to a person, not swappable into a codebase.
choose when you want a CLI for LandingAI's hosted extraction API rather than running a local model; seed supports local processing.
The official CLI for Agentic Document Extraction (ADE) by LandingAI — parse documents and extract schema-shaped data from your terminal
These projects were analysed and named docstrange among their alternatives. The relationship is not symmetric — how docstrange rates them is a separate judgement, made when docstrange is analysed in its own right.
calls docstrange “when you need a lightweight Python API that also emits CSV/HTML for downstream apps.”
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
calls docstrange “if you need direct URL extraction alongside office and image formats in one library”
Get your documents ready for gen AI
calls docstrange “When you need multi-format extraction with both OCR and no-OCR paths.”
Convert PDF to markdown + JSON quickly with high accuracy
calls docstrange “Choose for a MIT-licensed Python library with a simpler API for general documents; it lacks PDF/UA accessibility tagging.”
PDF Parser for AI-ready data. Automate PDF accessibility. Open-source.
calls docstrange “Python converter with OCR and multiple output formats; choose when JSON/HTML are also needed.”
Convert Word, PowerPoint, Excel, OpenDocument, RTF, EPUB, CSV, and PDF to clean Markdown. Built in Rust, with Node.js and Python bindings.
calls docstrange “when you need a library that directly OCRs and converts PDFs/images/Office files to Markdown/JSON/CSV/HTML with MIT license.”
Convert documents to structured data effortlessly. Unstructured is open-source ETL solution for transforming complex documents into clean, structured formats for language models. Visit our website to learn more about our enterprise grade Platform product for production grade workflows, partitioning, enrichments, chunking and embedding.
calls docstrange “when you need multi-format output (MD/JSON/CSV/HTML) and are OK with a traditional OCR-based pipeline.”
OCR & Document Extraction using vision models
calls docstrange “when you need multi-format document extraction (not just PDF) with a configurable pipeline”
A Comprehensive Toolkit for High-Quality PDF Content Extraction