drop-in peer for high-quality document conversion to Markdown/JSON with MIT license and active development
Get your documents ready for gen AI
- 65.4k
- Python
- MIT
PDF content extraction / document parsing
PDF-Extract-Kit is a modular Python toolkit that integrates state-of-the-art models for layout detection, formula detection, formula recognition, and OCR to extract high-quality structured content from PDF documents.
Same problem, same approach. Swapping one for another is a config change, not a rewrite.
drop-in peer for high-quality document conversion to Markdown/JSON with MIT license and active development
Get your documents ready for gen AI
drop-in peer for PDF-to-Markdown conversion with high accuracy and Apache-2.0 license
Convert PDF to markdown + JSON quickly with high accuracy
Solves the same problem with a different architecture or at a different layer. Expect to rewrite the integration.
when you prefer a vision-LLM-based approach over a pipeline of specialized models
A high-quality PDF to Markdown tool based on large language model visual recognition. 一款基于大模型视觉识别的高质量PDF转Markdown工具
when you need multi-format document extraction (not just PDF) with a configurable pipeline
Extract and convert data from any document, images, pdfs, word doc, ppt or URL into multiple formats (Markdown, JSON, CSV, HTML) with intelligent structured data extraction and advanced OCR.
when you need a Rust-native, fast document parser with low overhead instead of Python
A fast, helpful, and open-source document parser
when you want OCR-free extraction with a benchmarking toolkit
An on-premises, OCR-free unstructured data extraction, markdown conversion and benchmarking toolkit. (https://idp-leaderboard.org/)
when you want a TypeScript vision-based OCR extraction library instead of Python
OCR & Document Extraction using vision models
An earlier or more limited way to do the job. Still the right call when you need something small, proven, or CPU-only.
when you focus specifically on scanned books and want a tailored conversion tool
PDF craft can convert PDF files into various other formats. This project will focus on processing PDF files of scanned books.
when you need only table extraction from PDFs with a lightweight library
Camelot: PDF Table Extraction for Humans
when you need simple raw text extraction from many formats with a Node.js module
node.js module for extracting text from html, pdf, doc, docx, xls, xlsx, csv, pptx, png, jpg, gif, rtf and more!
when you need a mature Java-based text/metadata extraction library for many file types
The Apache Tika toolkit detects and extracts metadata and text from over a thousand different file types (such as PPT, XLS, and PDF).
Does the same job, but ships as an app. Useful to a person, not swappable into a codebase.
when you want a deployable API for document extraction with anonymization, not an embeddable library
Document (PDF, Word, PPTX ...) extraction and parse API using state of the art modern OCRs + Ollama supported models. Anonymize documents. Remove PII. Convert any document or picture to structured JSON or Markdown
when you need a ready-to-deploy Docling backend server rather than integrating the toolkit
Easily deployable and scalable backend server that efficiently converts various document formats (pdf, docx, pptx, html, images, etc) into Markdown. With support for both CPU and GPU processing, it is Ideal for large-scale workflows, it offers text/table extraction, OCR, and batch processing with sync/async endpoints.
when you want a CLI for agentic schema-shaped extraction rather than a library
The official CLI for Agentic Document Extraction (ADE) by LandingAI — parse documents and extract schema-shaped data from your terminal
when you want a self-hosted document parsing service with a UI/API
Transforms PDF, Documents and Images into Enriched Structured Data
These projects were analysed and named PDF-Extract-Kit among their alternatives. The relationship is not symmetric — how PDF-Extract-Kit rates them is a separate judgement, made when PDF-Extract-Kit is analysed in its own right.
calls PDF-Extract-Kit “when you want an AGPL-licensed modular PDF extraction toolkit with high-quality layout outputs.”
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
calls PDF-Extract-Kit “Choose if you want a modular, high-quality extraction toolkit (AGPL-3.0) for custom pipelines.”
Toolkit for linearizing PDFs for LLM datasets/training
calls PDF-Extract-Kit “Choose for a comprehensive Python toolkit for PDF content extraction including OCR, layout, formula, and table.”
GLM-OCR: Accurate × Fast × Comprehensive
calls PDF-Extract-Kit “research-oriented PDF extraction toolkit under AGPL-3.0, similar scope but copyleft license.”
Transforms PDF, Documents and Images into Enriched Structured Data
calls PDF-Extract-Kit “Comprehensive PDF content extraction with text, formula, and table recognition, a strong peer to OpenOCR for document parsing.”
OpenOCR: An Open-Source Toolkit for General-OCR Research and Applications, integrates a unified training and evaluation benchmark, commercial-grade OCR and Document Parsing systems, and faithful reproductions of the core implementations from a wide range of academic papers.
calls PDF-Extract-Kit “choose when you want to assemble your own PDF extraction modules (layout, formula, table) under AGPL; seed is an all-in-one tool.”
Extract and convert data from any document, images, pdfs, word doc, ppt or URL into multiple formats (Markdown, JSON, CSV, HTML) with intelligent structured data extraction and advanced OCR.
calls PDF-Extract-Kit “if you need a PDF-only extraction toolkit and can accept AGPL licensing”
Get your documents ready for gen AI
calls PDF-Extract-Kit “When you only need component extraction (text, formulas, tables) and want to assemble your own pipeline.”
Convert PDF to markdown + JSON quickly with high accuracy