Use instead when you want OCR directly via vision models rather than Tesseract plus LLM correction.
OCR & Document Extraction using vision models
- 12.3k
- TypeScript
- MIT
OCR correction / scanned PDF processing
Enhances Tesseract OCR output using LLMs for error correction, smart chunking, and markdown formatting of scanned PDFs
Solves the same problem with a different architecture or at a different layer. Expect to rewrite the integration.
Use instead when you want OCR directly via vision models rather than Tesseract plus LLM correction.
OCR & Document Extraction using vision models
Choose if you need an end-to-end OCR model without a separate LLM correction step.
GLM-OCR: Accurate × Fast × Comprehensive
Use when you want a dedicated OCR model for complex layouts rather than LLM-corrected Tesseract.
OCR model that handles complex tables, forms, handwriting with full layout.
Same document image parsing goal with a different model architecture.
The official repo for “Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting”, ACL, 2025.
Choose for an all-in-one OCR toolkit with more languages and layout awareness than Tesseract alone.
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
Use when you need a full document-to-markdown pipeline with layout and table recognition, not just OCR correction.
Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
Complete document preparation toolkit for gen AI, similar end goal.
Get your documents ready for gen AI
Alternative OCR model with optical compression, for those needing a different OCR backend.
Contexts Optical Compression
Use for ETL-style document processing into structured data for LLMs.
Convert documents to structured data effortlessly. Unstructured is open-source ETL solution for transforming complex documents into clean, structured formats for language models. Visit our website to learn more about our enterprise grade Platform product for production grade workflows, partitioning, enrichments, chunking and embedding.
Alternative document-to-structured-data tool with OCR support.
Transforms PDF, Documents and Images into Enriched Structured Data
Use for scanned PDF conversion to other formats, similar processing approach.
PDF craft can convert PDF files into various other formats. This project will focus on processing PDF files of scanned books.
Use for converting complex documents to RAG-ready data with vision models.
Vision infrastructure to turn complex documents into RAG/LLM-ready data
Use when you want an API service with OCR plus LLM, supporting more formats than the seed.
Document (PDF, Word, PPTX ...) extraction and parse API using state of the art modern OCRs + Ollama supported models. Anonymize documents. Remove PII. Convert any document or picture to structured JSON or Markdown
Does the same job, but ships as an app. Useful to a person, not swappable into a codebase.
End-user desktop OCR app, not embeddable in a codebase.
OCR software, free and offline. 开源、免费的离线OCR软件。支持截屏/批量导入图片,PDF文档识别,排除水印/页眉页脚,扫描/生成二维码。内置多国语言库。
End-user translation and OCR desktop app.
🌈一个跨平台的划词翻译和OCR软件 | A cross-platform software for text translation and recognition.
End-user Windows translation/OCR tool.
A ready-to-go translation ocr tool developed with WPF/WPF 开发的一款即用即走的翻译、OCR工具
End-user screenshot/OCR/search tool.
截屏 离线OCR 搜索翻译 以图搜图 贴图 录屏 万向滚动截屏 屏幕翻译 Screenshot Offline OCR Search Translate Search for picture Paste the picture on the screen Screen recorder Omnidirectional scrolling screenshot Screen translator 支持Windows Linux macOS
End-user screenshot/OCR/translation app.
🚀 Screenshots, word marking, OCR, AI, translation software || 截图、划词、文字识别、AI、翻译软件
These projects were analysed and named llm_aided_ocr among their alternatives. The relationship is not symmetric — how llm_aided_ocr rates them is a separate judgement, made when llm_aided_ocr is analysed in its own right.
calls llm_aided_ocr “choose when you want to improve existing Tesseract output with LLM error correction and markdown formatting.”
Contexts Optical Compression
calls llm_aided_ocr “choose for a composable Tesseract+LLM pipeline with explicit post-processing control.”
OCR model that handles complex tables, forms, handwriting with full layout.
calls llm_aided_ocr “Combines Tesseract OCR with LLM correction and formatting; choose when you want to use local/API LLMs for post-processing instead of a unified vision model.”
PDF craft can convert PDF files into various other formats. This project will focus on processing PDF files of scanned books.
calls llm_aided_ocr “Tesseract + LLM pipeline; choose for batch scanned-PDF to Markdown conversion”
PandaOCR - 多功能OCR图文识别+翻译+朗读+弹窗+公式+表格+图床+搜图+二维码