Skip to content
All tags

#document-parsing

14 posts

Commercial Document Parsing APIs Compared: Specialized Parsers, General VLMs, and the Big Three Clouds

Three routes to commercial document parsing: specialized parsers (Cohere Parse at $1.50/k pages, LlamaParse Agentic Plus at 90.2% on ParseBench), Big Three cloud prebuilts (Azure/Google/AWS for structured field extraction), and general-purpose VLMs (Fable 5.1 scores 78.92 on ParseBench and crushes specialized parsers on charts, but costs 3–16× more and hallucinates). At 100K pages/month, plain OCR runs ~$150 across providers; add tables and AWS jumps to $1,500, Claude Sonnet 5 to $900. The first question isn't 'which is most accurate' — it's 'do you need transcription or comprehension?'

Groundlane Series Part 6: The Document Toolkit — effort Dial, Field-Aware Chunking, and Confidence Routing

document_parse gains an effort parameter (fast/standard/deep) unifying anydoc WASM, OCR.space, and Docling VLM into a single dial; document_chunk's fieldAware mode extracts field names from table headers and metadata to solve RAG attribute conflation; document_smart_parse now returns a confidence score so agents decide whether to upgrade.

techdeep-dive

Groundlane Series Part 7: Selector Healing, Retrieval Test, and Quality Benchmarks

Groundlane v0.1.0 adds four quality mechanisms: web_extract selector healing (three-layer deterministic fallback, no LLM), corpus_retrieval_test (RAG recall verification inspired by RAGFlow), a document benchmark CLI (character-level F1), and a search benchmark framework (ground truth corpus with multi-provider comparison). Tool count goes from 54 to 55.

Marker: Datalab's Open-Source Pipeline Parser — Faster, CPU-Ready, More Accurate

Marker (Datalab open source, Apache 2.0 license, v2.0.0 released 2026-07-20, 39.5k stars) is a pipeline-style document parsing standard library that outputs Markdown and JSON, supports optional LLM boost (`--use_llm`, default `gemini-3.5-flash`), custom formatting logic, table/formula/inline-math/link/reference/code formatting, image extraction and preservation, header/footer removal, and runs on pure CPU, GPU, or MPS (`Apple Silicon`). Unlike [Docling](/posts/tech/2026-09-06-docling-document-parsing) (structured JSON core, dedicated XML exports, pure MIT), Marker centers on Markdown/JSON under `Apache 2.0` with a separate model-weight license (`AI Pubs Open Rail-M`, $5M commercial threshold); unlike [MinerU](/posts/tech/2026-09-05-mineru-ocr-doc-parsing) (custom agreement with MAU/revenue thresholds + attribution obligations), Marker offers a simpler licensing story (`Apache 2.0` code) but requires accepting a separate model-weight license for weights. The series framework ([three-layer model](/posts/ai/2026-08-06-document-parsing-three-layers)) positions all three (`MinerU`, `Docling`, `Marker`) as pipeline-based parsing-layer options with distinct licensing, output, and speed trade-offs.

Docling: IBM's Open-Source, MIT-Licensed Document Parsing Standard Library Built on Structured JSON

Docling (IBM Research Zurich, now governed by the Linux Foundation AI & Data, MIT license, v2.100.0 released 2026-06-09) is an open-source document parsing standard library with structured JSON (DoclingDocument) as its core output. It supports PDF, DOCX, PPTX, XLSX, HTML, EPUB, Apple Pages, video (MP4/AVI/MOV with ASR transcription and keyframes), audio (WAV/MP3), email (EML/MSG), ODF, and XBRL financial reports, with a swappable-stage pipeline parser (pure CPU or GPU-accelerated), VlmPipeline option (GraniteDocling 258M VLM), MCP server and API server (docling-serve), and native integrations with LangChain, LlamaIndex, Crew AI, and Haystack.

MinerU:從 PDF 到 Markdown 的開源文件解析引擎,為 RAG 與 LLM 訓練而生

MinerU(OpenDataLab / opendatalab/MinerU)是一個開源文件解析引擎,把 PDF、圖像、DOCX、PPTX、XLSX 轉成結構化 Markdown 與 JSON,內建公式識別、表格提取與 109 語言 OCR,起源於 InternLM 前訓練階段,適合 RAG 前處理與知識庫建構。

Agentic Parsing: Letting Agents Decide How to Parse Documents

Traditional document parsing runs a fixed pipeline regardless of input, but contracts, financial reports, and technical manuals each need different strategies. Agentic Parsing lets LLM agents observe a document and dynamically choose tools — AgenticOCR parses only the regions that matter (70%+ visual token savings), and ParseBench shows even the best method scores only 84.9% across 2,000 enterprise pages. No silver bullet.

AI Agent GitHub Digest — 2026-09-03

NousResearch/hermes-agent keeps climbing (239,994 stars) on a self-improving learning loop that remembers how to use your tools and who you are across sessions. pacifio/atlas gained 895 stars in a day by giving multiple coding agents shared, traceable version control — every commit links back to the session that made it. blader/humanizer strips the AI tell from writing using 35 patterns, without inventing facts. On the document side, firecrawl/pdf-inspector decides in under 50ms whether a PDF needs OCR, and superlinked/sie folds every model an agent needs into one self-hosted inference cluster. On the framework side, AG2 v1.0.3 ports fully to MCP 2.0 (a breaking change) and adds TealTigerMiddleware, a deterministic, non-LLM prompt-injection guard.

aideep-dive

RAGFlow Deep Dive: From Document Parsing and Chunk Review to Cited Answers

RAGFlow puts document parsing, human chunk review, retrieval tests, chat, and citations in one platform; it fits layout-heavy PDFs and tables, but carries more deployment weight and platform state than a Python library.

Scanned PDF Benchmark: How Did 10 Parsers Handle Graduate Entrance Exams?

I tested 10 open-source PDF parsing tools on four scanned NTU graduate entrance exams. VLM-based tools—Firecrawl, MinerU 3.4, and Marker v2—overwhelmingly beat conventional OCR on formulas and code, but installation was the real barrier: MinerU's old package name creates dependency hell, Marker's first model download takes 10 minutes, and PaddleOCR needs a separate engine. In practice, use RapidOCR for screening and MinerU or Firecrawl for close inspection.

The Parsing Layer: When Structure Must Be Inferred — and Licensing Is the Real Selection Axis

Scans and complex layouts leave you no choice but to infer structure with a model. But the technical gap between MinerU, Marker, and Docling is far smaller than the licensing gap — MinerU needs a separate license past $20M monthly revenue, Marker's model weights need payment past a funding threshold, and only Docling is cleanly MIT. Read the LICENSE before the benchmark.

The Three-Layer Ladder of Document Parsing: Pick the Layer Before the Tool

The most common mistake in feeding documents to an LLM isn't picking the wrong tool — it's picking the wrong layer. Structure already in the file goes to the conversion layer (milliseconds); text without structure goes to extraction; only inferred structure needs parsing. anydoc's 4.7ms against Docling's 513.6ms is a 109× gap, and most people jump straight to the most expensive layer.

The Deterministic Extraction Layer: Solve 80% of Your PDFs With No Model At All

Digital-native PDFs already contain readable text — what's missing is structure, and heuristics can recover it. PyMuPDF, pdfplumber, pypdf, and Tika do this with zero GPU and zero inference cost. The biggest selection trap isn't accuracy; it's PyMuPDF's AGPL-3.0 license.

aideep-dive

Auto-Embedding on File Upload Is a Bad Default: A Survey of Adaptive / Agentic RAG and Agentic Parsing

Making 'chunk and embed every uploaded file automatically' the default behavior means making a decision for the LLM that it could have made itself. From Self-RAG (2310.11511) and Adaptive-RAG (2403.14403) to AgenticOCR (2602.24134), the academic trajectory is pushing three layers of decision-making -- whether to retrieve, whether to parse, and how to chunk -- from the ingestion pipeline back to the agent at conversation time.