AltHub

oomol-lab/pdf-craft alternatives

scanned PDF to Markdown/EPUB conversion

Converts scanned book PDFs to Markdown/EPUB using DeepSeek OCR locally, preserving tables, formulas, and footnotes.

6.2kPythonMITactive · last push 1d agoGitHub

Drop-in peers

Same problem, same approach. Swapping one for another is a config change, not a rewrite.

getomni-ai/zerox

Vision-model OCR to Markdown; choose when you need a model-agnostic vision OCR API or prefer OpenAI-compatible backends over DeepSeek.

OCR & Document Extraction using vision models

12.3k
TypeScript
MIT
slowing1.3y ago
bytedance/Dolphin

Vision-language model for document image parsing; choose when you want to use this specific model directly for Markdown generation.

The official repo for “Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting”, ACL, 2025.

9.1k
Python
NOASSERTION
active5mo ago
NanoNets/docext

OCR-free markdown conversion using vision models; a direct peer when you prefer an on-premises vision-model-based toolkit.

An on-premises, OCR-free unstructured data extraction, markdown conversion and benchmarking toolkit. (https://idp-leaderboard.org/)

2.1k
Python
Apache-2.0
active5mo ago

Same job, different approach

Solves the same problem with a different architecture or at a different layer. Expect to rewrite the integration.

run-llama/liteparse

Fast Rust document parser for text-based PDFs; choose when you don't need OCR and want a lightweight library for Markdown extraction.

A fast, helpful, and open-source document parser

12.2k
Rust
Apache-2.0
active2d ago
CatchTheTornado/text-extract-api

API service combining classic OCR with LLM post-processing; pick when you want a self-hosted extraction API rather than a Python library.

Document (PDF, Word, PPTX ...) extraction and parse API using state of the art modern OCRs + Ollama supported models. Anonymize documents. Remove PII. Convert any document or picture to structured JSON or Markdown

3.2k
Python
MIT
slowing9mo ago
axa-group/Parsr

Extracts structured data from PDFs/images using traditional document analysis; choose when you need JSON output rather than Markdown/EPUB.

Transforms PDF, Documents and Images into Enriched Structured Data

6.2k
JavaScript
Apache-2.0
active5mo ago
docling-project/docling

Pipeline using layout models and external OCR engines; a flexible open-source converter without relying on a commercial vision API.

Get your documents ready for gen AI

65.4k
Python
MIT
active2d ago
opendataloader-project/opendataloader-pdf

Java-based PDF parser for AI-ready data; a JVM-native alternative for text-based PDFs, but not specialized for scanned books.

PDF Parser for AI-ready data. Automate PDF accessibility. Open-source.

28.6k
Java
Apache-2.0
active2d ago
Dicklesworthstone/llm_aided_ocr

Combines Tesseract OCR with LLM correction and formatting; choose when you want to use local/API LLMs for post-processing instead of a unified vision model.

Enhances Tesseract OCR output using LLMs (local or API) for error correction, smart chunking, and markdown formatting of scanned PDFs

3.0k
Python
NOASSERTION
active20d ago
opendatalab/MinerU

Mature pipeline for converting complex PDFs to Markdown/JSON using layout detection and OCR; a strong but internally different alternative.

Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.

78.2k
Python
NOASSERTION
active4d ago
Unstructured-IO/unstructured

ETL for document structuring with multiple OCR backends; choose when you need to handle many document types beyond just PDFs.

Convert documents to structured data effortlessly. Unstructured is open-source ETL solution for transforming complex documents into clean, structured formats for language models. Visit our website to learn more about our enterprise grade Platform product for production grade workflows, partitioning, enrichments, chunking and embedding.

15.3k
HTML
Apache-2.0
active2d ago

Older generation, or narrower

An earlier or more limited way to do the job. Still the right call when you need something small, proven, or CPU-only.

ocrmypdf/OCRmyPDF

Adds a searchable text layer to scanned PDFs; pick when you only need searchable PDFs, not Markdown/EPUB conversion.

OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched

34.5k
Python
MPL-2.0
active1d ago

End-user tools

Does the same job, but ships as an app. Useful to a person, not swappable into a codebase.

ciur/papermerge

Document management system for archiving scanned PDFs, not a conversion library.

Open Source Document Management System for Digital Archives (Scanned Documents)

2.9k
Python
Apache-2.0
slowing9mo ago
PDFCraftTool/pdfcraft

Browser-based PDF toolkit with 90+ end-user features, not a swappable library for developers.

PDFCraft is a free, privacy-focused PDF toolkit that runs entirely in your browser. With 90+ professional tools, you can edit, convert, merge, split, and secure your PDF files without ever uploading them to a server.

8.3k
TypeScript
AGPL-3.0
active23d ago
ahrm/sioyek

PDF viewer optimized for research papers, not a conversion tool.

Sioyek is a PDF viewer with a focus on textbooks and research papers

9.8k
C
GPL-3.0
active5d ago

Listed as an alternative to

These projects were analysed and named pdf-craft among their alternatives. The relationship is not symmetric — how pdf-craft rates them is a separate judgement, made when pdf-craft is analysed in its own right.

ocrmypdf/OCRmyPDF

calls pdf-craftConverts scanned PDFs into other formats via OCR; choose when you want editable conversions rather than PDF/A searchable output.

OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched

34.5kPythonactive
getomni-ai/zerox

calls pdf-craftwhen you need a Python-based PDF converter focused on scanned books using a traditional OCR pipeline.

OCR & Document Extraction using vision models

12.3kTypeScriptslowing
Dicklesworthstone/llm_aided_ocr

calls pdf-craftUse for scanned PDF conversion to other formats, similar processing approach.

Enhances Tesseract OCR output using LLMs (local or API) for error correction, smart chunking, and markdown formatting of scanned PDFs

3.0kPythonactive
allenai/olmocr

calls pdf-craftChoose when your corpus is mainly scanned books and you need a simpler, MIT-licensed tool.

Toolkit for linearizing PDFs for LLM datasets/training

19.4kPythonactive
alam00000/bentopdf

calls pdf-craftSpecialized PDF-to-other-format converter focused on scanned books; choose it for scanned-document conversion rather than general PDF manipulation.

The Privacy First PDF Toolkit

14.7kJavaScriptactive
datalab-to/chandra

calls pdf-craftchoose for converting scanned books with a simpler, MIT-licensed tool.

OCR model that handles complex tables, forms, handwriting with full layout.

12.1kPythonactive
opendatalab/PDF-Extract-Kit

calls pdf-craftwhen you focus specifically on scanned books and want a tailored conversion tool

A Comprehensive Toolkit for High-Quality PDF Content Extraction

10.0kPythondormant
axa-group/Parsr

calls pdf-craftnarrower, focuses on scanned-book PDF conversion to formats rather than full document structure extraction.

Transforms PDF, Documents and Images into Enriched Structured Data

6.2kJavaScriptactive
NanoNets/docstrange

calls pdf-craftchoose when your workload is scanned book PDFs and you need specialized book layout handling; seed covers more formats but with less focus.

Extract and convert data from any document, images, pdfs, word doc, ppt or URL into multiple formats (Markdown, JSON, CSV, HTML) with intelligent structured data extraction and advanced OCR.

1.5kPythonslowing
opendataloader-project/opendataloader-pdf

calls pdf-craftChoose for a ready-made app that converts scanned book PDFs to other formats; you get a UI/CLI rather than an API.

PDF Parser for AI-ready data. Automate PDF accessibility. Open-source.

28.6kJavaactive
zai-org/GLM-OCR

calls pdf-craftUse as an end-user PDF conversion tool for scanned books.

GLM-OCR: Accurate × Fast × Comprehensive

7.3kPythonactive