PDF to Structured JSON for RAG and LLM Pipelines
Extracting PDFs correctly prevents silent corruption that degrades RAG and LLM retrieval quality.

PDF extraction is an engineering problem with its own failure modes, and getting it wrong corrupts everything built on top of it. PDFs were built to look right on a screen or a printed page, not to hand their structure to a machine. A table looks like a table to a person because of where the lines and spacing sit, while the file itself contains no marker saying "this is a table. Every parser has to guess that structure back, and a lot of parsers guess wrong.
The gap between extracting something and extracting something a language model can actually use is wider than most tutorials let on. A 2024 survey from Peking University and Shanghai AI Lab, cited in a review of PDF parsers for AI workflows, draws the line clearly between basic OCR, which pulls out words. A real document parser knows a heading is a heading, knows a table has rows and columns that relate to each other, and knows a multi-column layout has to be read in a specific order, not straight across the page.
A RAG or LLM pipeline takes on whatever structure the extraction layer gives it, good or bad. If a chunk comes out of the parser corrupted, that chunk gets indexed exactly as it is, errors and all. Whether a RAG system receives reliable input or corrupted chunks depends on whether the parser returns only raw text or preserves structure such as headings, table rows and columns, and reading order. Invofox validates extraction against a defined schema and runs layout detection before any language model processes the page, so the model works from structured, traceable data instead of a reconstructed guess.
Silent extraction failures and retrieval index corruption
The extraction failures that do real damage are not crashes. The dangerous failures are the ones where the parser returns output that looks fine and isn't, and the retrieval system indexes it without complaint.
The 2026 PureDocBench study names several of these failure types directly. A third is a technical symbol recognition error, where a unit mutates during extraction, for instance nH turning into mH. That one character swap is a million-fold error in the actual value, and in a safety-critical specification document, a mistake like that can change what the document is understood to mean.
Layout errors compound once they start. A single misread column boundary can corrupt an entire section's worth of retrieved content, not just the one line where the error started. Headers, footers, and figure captions bleed into the body text, diluting whatever signal the embedding model was supposed to pick up from that chunk.
What makes this worse operationally is that none of it announces itself. No alarm goes off. Catching these failures before indexing requires instrumentation: per-field confidence scores tied to source locations, not inferred values, combined with structural sanity checks at ingestion. Invofox grounds each extracted field to its source pixels and tracks field-level accuracy, which makes these silent failures observable instead of letting them erode retrieval quality unnoticed.
Why aggregate accuracy benchmarks do not predict production behavior
A high leaderboard rank does not mean a model reproduces complex documents correctly. The score and the behavior are not the same thing. The same study found that agreement between models carries its own signal: when models disagree with each other, that can point to hallucination, and when models agree with each other but the accuracy is still low, that can mean the reference data itself has errors in it.
The 2026 ParseBench framework breaks this down into five extraction dimensions that a document pipeline must satisfy to hold up in production. Tables need structural fidelity across merged cells, hierarchical headers, and continuation across page breaks, because a single shifted header silently pulls the wrong value into a field and gives no indication anything went wrong. Charts need exact data-point values pulled out, not a paragraph summarizing what the chart shows, because an agent reasoning over a chart needs the number, not a description of the number. Content faithfulness covers omissions, hallucinations, and reading-order mistakes, since dropped or invented content means an agent is acting on context that was never actually in the document. Semantic formatting, things like strikethrough, superscript, subscript, bold text, and hyperlinks, carries real meaning in financial and legal documents, and losing it changes what the document says. Visual grounding rounds out the fifth dimension, tying every extracted value back to where it actually sits on the page. Rigorous evaluation treats each of these as its own measurement: every extracted leaf field gets a correctness label matched to its schema type, exact match, numeric match, or semantic match, with omissions and hallucinations scored on their own rather than folded into one overall grade. That is a different way of measuring a parser than running it against a benchmark and reading off a single percentage, and it is the only kind of measurement that catches the failures described above before they reach a production index.
Grounding extracted fields to their source pixels as the primary defense against hallucination
A confidence score from a vision model is only useful if it reflects what the model actually read off the page. A model can report high confidence on a value it inferred from surrounding context rather than one it read directly, and the score alone gives no way to tell the difference. Confidence has to trace back to a specific location on the page: a bounding box, a page number, something concrete enough to check against the source document.
Testing on the CORD dataset in its hard evaluation regime found that only about half of all asserted fields came out correct, on the claude-sonnet-5 capture. But grounded fields, those tied to a traceable source location, were substantially more likely to be correct than ungrounded ones. That made grounding the strongest available signal for deciding whether to accept a field automatically or route it to human review, though the same pattern did not hold under other models such as haiku or qwen. Grounding's value as a signal is model-dependent and has to be validated per model rather than assumed to generalize.
Building this into a pipeline means every extracted field carries a source citation at minimum, a page number and a bounding box or character span pointing back to the exact pixels the value came from. A vendor that measures accuracy only at the document or page level hides the field-level variance that actually decides whether downstream retrieval works. Invofox's per-field service-level agreement and continuous measurement against individual fields, rather than aggregate scores, reveal which extraction dimensions fail in production and where, instead of letting a benchmark average mask real corruption in the pipeline.
Vision-language-model-based parsers carry their own specific risk here. The 2026 parser roundup notes that these models can hallucinate content in high-resolution, text-dense documents, inventing text that was never on the page. That is a worse failure than simply leaving a field blank, because a blank field is detectable and an invented one usually is not, at least not without the grounding check described above. PDFs can carry hidden text in their underlying text layer, invisible to a human reading the rendered page but passed straight to a language model if the pipeline does not separate visual rendering from that text layer. That gap opens a path for prompt injection, content planted in a document specifically to manipulate the model processing it, a mechanism distinct from hallucination.
Output modes for a production pipeline
A production pipeline needs two separate kinds of output, not one that tries to do both jobs. The first is full-document structured Markdown, built for RAG ingestion. The second is schema-defined JSON, built for typed field retrieval. Trying to force one format to serve both purposes makes both of them worse.
The Markdown path exists to feed whole documents into a vector store or into a model's context window. The Elastic and LlamaParse integration describes this directly: LlamaParse converts documents into Markdown meant to be read by a language model, keeping the full layout of the original document intact. That Markdown then goes through chunking and cleaning before the resulting pieces get embedded and stored in a vector database, ready for semantic search.
The JSON path solves a different problem. Instead of handing over the whole document, a developer defines a schema ahead of time specifying exactly which fields matter, and the extraction model returns those fields as validated JSON. In the same Elastic pipeline, this runs through LlamaParse Extract, built to pull values out of bar charts and other visual elements that a standard text parser would miss. That output lands as typed fields inside Elasticsearch and gets queried with a structured query language, a fundamentally different retrieval contract than a semantic vector search, built for pulling a specific number or date rather than a relevant passage.
Many production systems end up running both paths at once: structured JSON stored in a relational or document database, sitting alongside embeddings in a vector store, with the extraction layer responsible for producing both outputs accurately and keeping them consistent with each other. The scale this can reach shows up in MMORE, the open-source multimodal RAG pipeline from EPFL, ETH Zürich, and Harvard. It ingests more than fifteen file types, including text, tables, images, emails, audio, and video, processes all of it into one unified format, and supports hybrid dense-sparse retrieval along with both interactive API and batch RAG endpoints. That range of format support gives a sense of how much the extraction layer has to handle before either output mode above can even start doing its job.
Requirements for a reliable extraction layer
Neither output mode works if the extraction layer can't handle the structural complexity of the document in front of it. Most parsers handle one or two document types well and fall apart in their own particular way once pushed past that range, so the question for any pipeline is which document types it actually needs to cover, and whether the parser was built for those.
Tables are the clearest test case. Tables that break across a page boundary present a separate problem: the second half of the table has to get reassembled with the first, or it ends up treated as unrelated freestanding text with no connection to the headers that gave it meaning.
Multi-column layouts cause a different kind of damage. A parser that reads straight across the page instead of down one column and then the next interleaves two unrelated streams of text into a single chunk. That chunk makes no sense as a unit, it embeds as a kind of semantic non-sequitur, and whatever gets retrieved from it reads like nonsense to anyone looking at the query result.
Scanned and image-based PDFs need an entirely different processing stack than text-based ones, and this is where aggregate accuracy numbers mislead the most. Charts and other visual elements raise a related issue: the values embedded in a bar chart, a scatter plot, or an infographic exist only as pixels on the page. A parser without visual element extraction returns nothing for those fields, and an agent reasoning over the document has no way of knowing a gap is even there. The nH-to-mH unit mutation described earlier falls into this same category of risk, a formula and symbol fidelity failure, where a parser built for scientific and technical documents needs to preserve mathematical expressions in a form a machine can still read correctly, rather than returning garbled characters or dropping the formula.
The build-versus-buy calculus when extraction quality is the constraint
The biggest mistake in planning an extraction pipeline is treating parsing as something to configure once and leave alone. Documents arriving three months from now will not match the layouts and vendor formats indexed last quarter, and every new format or layout shift means another manual update to templates or configuration. At scale, that maintenance cost grows in direct proportion to how many vendors a pipeline ingests from: an organization pulling documents from a hundred different vendors ends up maintaining something close to a hundred different templates, and each one needs attention whenever the vendor changes its layout.
Building a custom pipeline makes sense in specific circumstances. Self-hosted, open-source tools such as IBM's Docling offer high-fidelity table extraction and reading order detection with no dependency on an outside vendor, which fits well when privacy, cost control, or offline operation is the deciding constraint. But standing up and running that kind of infrastructure is itself a substantial engineering project, not a weekend setup, and it carries the same honest limitation every custom pipeline carries: every component that handles an edge case correctly in testing will eventually meet a variant it has never seen before in production, and the system needs to say "unable to process this" explicitly rather than guess and move on.
That is the actual tradeoff between building and buying when extraction quality is the limiting factor. Building buys control and independence from a vendor, at the cost of an ongoing engineering commitment that grows with every new document format. Buying a service built specifically around field-level accuracy, grounding, and continuous measurement against real documents shifts that maintenance burden onto a vendor whose entire job is keeping up with exactly the kind of format drift that breaks custom pipelines over time. The choice comes down to which side of that maintenance cost a team is actually equipped to carry.


