choose as a drop-in parser for PDF/Office to Markdown/JSON with strong OCR, but you'll handle PII separately.
Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
- 78.2k
- Python
- NOASSERTION
document extraction / OCR / PII redaction
FastAPI-based document extraction API that converts PDFs, images, and Office files into Markdown or JSON using OCR + Ollama LLMs, with PII removal.
Same problem, same approach. Swapping one for another is a config change, not a rewrite.
choose as a drop-in parser for PDF/Office to Markdown/JSON with strong OCR, but you'll handle PII separately.
Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
Solves the same problem with a different architecture or at a different layer. Expect to rewrite the integration.
choose for a Rust-native document parser when you only need digital text extraction, not OCR or PII removal.
A fast, helpful, and open-source document parser
choose for fast conversion of digital Office/PDF docs to Markdown in Rust with Python/Node bindings, no OCR or PII.
Convert Word, PowerPoint, Excel, OpenDocument, RTF, EPUB, CSV, and PDF to clean Markdown. Built in Rust, with Node.js and Python bindings.
choose for a Python library converting office/PDF docs to Markdown, but with no PII removal and a simpler pipeline.
Python tool for converting files and office documents to Markdown.
An earlier or more limited way to do the job. Still the right call when you need something small, proven, or CPU-only.
choose when you only need text-level PII redaction as a TypeScript library, not document parsing.
Remove personally identifiable information from text.
Java library/service for redacting PII/PHI in text, narrower than full extraction.
Philter redacts sensitive information such as PII and PHI in text.
Python library for context-aware PII detection and rewriting, narrower than full document parsing.
๐ต๏ธ NeMo Anonymizer: Detect and protect PII through context-aware replacement and rewriting
Does the same job, but ships as an app. Useful to a person, not swappable into a codebase.
desktop app for PII removal in documents, not a library you can embed.
Desktop App with Built-In LLM for Removing Personal Identifiable Information in Documents
deploy as a proxy to strip PII from OpenAI API calls, not an embeddable library.
SanitAI is a drop-in proxy for OpenAI's API to detect and remove PII data.
CLI/MCP tool for Claude and command-line PII redaction, not an embeddable library.
PII anonymization (MCP + Skill for Claude and CLI for the rest)
native macOS app for PII removal using Vision OCR, end-user tool.
Native macOS PII removal software with on device Vision OCR and OpenAI privacy filter model
These projects were analysed and named text-extract-api among their alternatives. The relationship is not symmetric โ how text-extract-api rates them is a separate judgement, made when text-extract-api is analysed in its own right.
calls text-extract-api โAPI service combining classic OCR with LLM post-processing; pick when you want a self-hosted extraction API rather than a Python library.โ
PDF craft can convert PDF files into various other formats. This project will focus on processing PDF files of scanned books.
calls text-extract-api โUse when you want an API service with OCR plus LLM, supporting more formats than the seed.โ
Enhances Tesseract OCR output using LLMs (local or API) for error correction, smart chunking, and markdown formatting of scanned PDFs
calls text-extract-api โA deployable extraction API service, not importable as a library.โ
Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
calls text-extract-api โChoose if you want to deploy a self-hosted API for document extraction with PII removal; it's a service, not a library.โ
PDF Parser for AI-ready data. Automate PDF accessibility. Open-source.
calls text-extract-api โwhen you want to run a self-hosted extraction API that returns JSON, rather than integrating a library.โ
OCR & Document Extraction using vision models
calls text-extract-api โstandalone extraction API wrapping multiple OCR engines; not a model import.โ
OCR model that handles complex tables, forms, handwriting with full layout.
calls text-extract-api โwhen you want a deployable API for document extraction with anonymization, not an embeddable libraryโ
A Comprehensive Toolkit for High-Quality PDF Content Extraction