Choose for Apache-2.0 parser with broader input support (DOCX/PPTX) and a goal of no-loss parsing.
File Parser optimised for LLM Ingestion with no loss 🧠 Parse PDFs, Docx, PPTx in a format that is ideal for LLMs.
- 7.4k
- Python
- Apache-2.0
PDF/document to text conversion
Toolkit for converting PDFs and other image-based documents into clean, readable plain text/Markdown for LLM datasets and training.
Same problem, same approach. Swapping one for another is a config change, not a rewrite.
Choose for Apache-2.0 parser with broader input support (DOCX/PPTX) and a goal of no-loss parsing.
File Parser optimised for LLM Ingestion with no loss 🧠 Parse PDFs, Docx, PPTx in a format that is ideal for LLMs.
Choose if you want a widely used, Apache-2.0 converter with high accuracy and active development.
Convert PDF to markdown + JSON quickly with high accuracy
Choose for a very active, feature-rich option that converts PDFs and Office docs to markdown/JSON.
Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
Choose if you want an MIT-licensed, highly configurable document conversion library with IBM backing.
Get your documents ready for gen AI
Choose if you want a modular, high-quality extraction toolkit (AGPL-3.0) for custom pipelines.
A Comprehensive Toolkit for High-Quality PDF Content Extraction
Choose if you work in the JVM/Java ecosystem and need a PDF parser tailored for AI-ready data.
PDF Parser for AI-ready data. Automate PDF accessibility. Open-source.
Solves the same problem with a different architecture or at a different layer. Expect to rewrite the integration.
Choose if you need multi-language OCR and a traditional pipeline that can be customized for many document types.
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
Choose if you want a JavaScript-based document-to-structured-data pipeline with a REST API.
Transforms PDF, Documents and Images into Enriched Structured Data
An earlier or more limited way to do the job. Still the right call when you need something small, proven, or CPU-only.
Choose when your corpus is mainly scanned books and you need a simpler, MIT-licensed tool.
PDF craft can convert PDF files into various other formats. This project will focus on processing PDF files of scanned books.
These projects were analysed and named olmocr among their alternatives. The relationship is not symmetric — how olmocr rates them is a separate judgement, made when olmocr is analysed in its own right.
calls olmocr “choose when you need to batch-convert PDFs to plain text for LLM training corpora; seed outputs Markdown/JSON and is not a training-data linearizer.”
Extract and convert data from any document, images, pdfs, word doc, ppt or URL into multiple formats (Markdown, JSON, CSV, HTML) with intelligent structured data extraction and advanced OCR.