Vision-model OCR to Markdown; choose when you need a model-agnostic vision OCR API or prefer OpenAI-compatible backends over DeepSeek.
OCR & Document Extraction using vision models
- 12.3k
- TypeScript
- MIT
scanned PDF to Markdown/EPUB conversion
Converts scanned book PDFs to Markdown/EPUB using DeepSeek OCR locally, preserving tables, formulas, and footnotes.
Same problem, same approach. Swapping one for another is a config change, not a rewrite.
Vision-model OCR to Markdown; choose when you need a model-agnostic vision OCR API or prefer OpenAI-compatible backends over DeepSeek.
OCR & Document Extraction using vision models
Vision-language model for document image parsing; choose when you want to use this specific model directly for Markdown generation.
The official repo for “Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting”, ACL, 2025.
OCR-free markdown conversion using vision models; a direct peer when you prefer an on-premises vision-model-based toolkit.
An on-premises, OCR-free unstructured data extraction, markdown conversion and benchmarking toolkit. (https://idp-leaderboard.org/)
Solves the same problem with a different architecture or at a different layer. Expect to rewrite the integration.
Fast Rust document parser for text-based PDFs; choose when you don't need OCR and want a lightweight library for Markdown extraction.
A fast, helpful, and open-source document parser
API service combining classic OCR with LLM post-processing; pick when you want a self-hosted extraction API rather than a Python library.
Document (PDF, Word, PPTX ...) extraction and parse API using state of the art modern OCRs + Ollama supported models. Anonymize documents. Remove PII. Convert any document or picture to structured JSON or Markdown
Extracts structured data from PDFs/images using traditional document analysis; choose when you need JSON output rather than Markdown/EPUB.
Transforms PDF, Documents and Images into Enriched Structured Data
Pipeline using layout models and external OCR engines; a flexible open-source converter without relying on a commercial vision API.
Get your documents ready for gen AI
Java-based PDF parser for AI-ready data; a JVM-native alternative for text-based PDFs, but not specialized for scanned books.
PDF Parser for AI-ready data. Automate PDF accessibility. Open-source.
Combines Tesseract OCR with LLM correction and formatting; choose when you want to use local/API LLMs for post-processing instead of a unified vision model.
Enhances Tesseract OCR output using LLMs (local or API) for error correction, smart chunking, and markdown formatting of scanned PDFs
Mature pipeline for converting complex PDFs to Markdown/JSON using layout detection and OCR; a strong but internally different alternative.
Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
ETL for document structuring with multiple OCR backends; choose when you need to handle many document types beyond just PDFs.
Convert documents to structured data effortlessly. Unstructured is open-source ETL solution for transforming complex documents into clean, structured formats for language models. Visit our website to learn more about our enterprise grade Platform product for production grade workflows, partitioning, enrichments, chunking and embedding.
An earlier or more limited way to do the job. Still the right call when you need something small, proven, or CPU-only.
Adds a searchable text layer to scanned PDFs; pick when you only need searchable PDFs, not Markdown/EPUB conversion.
OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched
Does the same job, but ships as an app. Useful to a person, not swappable into a codebase.
Document management system for archiving scanned PDFs, not a conversion library.
Open Source Document Management System for Digital Archives (Scanned Documents)
Browser-based PDF toolkit with 90+ end-user features, not a swappable library for developers.
PDFCraft is a free, privacy-focused PDF toolkit that runs entirely in your browser. With 90+ professional tools, you can edit, convert, merge, split, and secure your PDF files without ever uploading them to a server.
PDF viewer optimized for research papers, not a conversion tool.
Sioyek is a PDF viewer with a focus on textbooks and research papers
These projects were analysed and named pdf-craft among their alternatives. The relationship is not symmetric — how pdf-craft rates them is a separate judgement, made when pdf-craft is analysed in its own right.
calls pdf-craft “Converts scanned PDFs into other formats via OCR; choose when you want editable conversions rather than PDF/A searchable output.”
OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched
calls pdf-craft “when you need a Python-based PDF converter focused on scanned books using a traditional OCR pipeline.”
OCR & Document Extraction using vision models
calls pdf-craft “Use for scanned PDF conversion to other formats, similar processing approach.”
Enhances Tesseract OCR output using LLMs (local or API) for error correction, smart chunking, and markdown formatting of scanned PDFs
calls pdf-craft “Choose when your corpus is mainly scanned books and you need a simpler, MIT-licensed tool.”
Toolkit for linearizing PDFs for LLM datasets/training
calls pdf-craft “Specialized PDF-to-other-format converter focused on scanned books; choose it for scanned-document conversion rather than general PDF manipulation.”
The Privacy First PDF Toolkit
calls pdf-craft “choose for converting scanned books with a simpler, MIT-licensed tool.”
OCR model that handles complex tables, forms, handwriting with full layout.
calls pdf-craft “when you focus specifically on scanned books and want a tailored conversion tool”
A Comprehensive Toolkit for High-Quality PDF Content Extraction
calls pdf-craft “narrower, focuses on scanned-book PDF conversion to formats rather than full document structure extraction.”
Transforms PDF, Documents and Images into Enriched Structured Data
calls pdf-craft “choose when your workload is scanned book PDFs and you need specialized book layout handling; seed covers more formats but with less focus.”
Extract and convert data from any document, images, pdfs, word doc, ppt or URL into multiple formats (Markdown, JSON, CSV, HTML) with intelligent structured data extraction and advanced OCR.
calls pdf-craft “Choose for a ready-made app that converts scanned book PDFs to other formats; you get a UI/CLI rather than an API.”
PDF Parser for AI-ready data. Automate PDF accessibility. Open-source.
calls pdf-craft “Use as an end-user PDF conversion tool for scanned books.”
GLM-OCR: Accurate × Fast × Comprehensive