Document AILong read

Document-Specific Foundation Models vs General LLMs for Extraction

Spatial structure matters far more than model size for document extraction.

Contributing Editor · · 10 min read
Cover illustration for “Document-Specific Foundation Models vs General LLMs for Extraction”
Document AI · October 2, 2026 · 10 min read · 2,298 words

A W-2 splits its boxes across two columns. A bill of lading often carries a consignee field that wraps across three or four lines. A commercial lease nests one table inside another. None of this is decoration. The position of a number on the page, next to a label, inside a box, is part of what that number means. The real question behind "general or document-specific model" is whether a system treats spatial structure as signal or noise.

General-purpose language models read documents as strings of words. They're built to reason over meaning and sequence, not over where a column starts or where a header ends. A June 2026 industry guide states the problem directly: "The document's meaning lives partly in its layout; a text-only model drops that signal entirely". It comes from how these models are built in the first place, not from a prompting mistake someone can engineer around: they turn a page into a sequence of tokens, and a sequence has no columns.

Document-specific architectures start from the opposite assumption. They treat layout zones, table cell boundaries, column headers, and multi-page context as inputs the model needs before it ever tries to pull out a value. The spatial map comes first; extraction follows it.

That one design choice, whether spatial structure counts as input or gets discarded, predicts with real precision where each kind of system will hold up and where it will break once it meets production volume and messy real-world paper. O The next section walks through what that breakage looks like, document type by document type.

R Layout blindness and repeatable production failure modes

LLM extraction errors on real-world documents don't scatter randomly across a file. They cluster on the specific features where position carries meaning that the surrounding words alone don't supply.

Take long documents first. A 200-page loan package or a multi-exhibit insurance submission, once converted to plain text, can run to enormous token counts. A A field defined on page 3 gets orphaned by the time it's referenced in a schedule on page 47; an exhibit number cited on one page loses its anchor to the exhibit itself; a table that splits across a chunk boundary produces two half-rows that get stitched into one mismatched value. None of this throws an error. No confidence score flags the loss, so the document moves downstream looking clean while carrying a wrong number.

Tables that cross page breaks cause a related but distinct problem. An LLM reading page-by-page text loses the thread between a row and its column once that row continues on the next page. A document-aware system that keeps context at the document level, not the page level, keeps the cell-to-column relationship intact across that boundary. Run a 40-page commercial insurance submission with embedded loss run tables through a text-sequence model, and rows that wrap across a page break can get assigned cell values that never belonged to them; dense tables lose their column alignment entirely.

Sparse or degraded input produces a third, more dangerous pattern. When a field is missing, worded ambiguously, or blurred by a scan artifact, a general LLM tends to generate a plausible-sounding value instead of returning null or flagging uncertainty. In mortgage underwriting or insurance intake, a confidently wrong field beats a blank one for all the wrong reasons: the blank gets caught by a review step, and the hallucinated value sails through validation because it looks like a real answer.

Two newer failure classes have shown up specifically in production vision-language-model pipelines across major API providers: repetition loops that chew through compute without producing output, and recitation filters that block extractions that were entirely legitimate. Both failures compound as they move through a pipeline. A character recognition error early in the process doesn't stay small. It turns into a much larger failure rate once that garbled text reaches the structured-extraction stage downstream.

Where general LLMs genuinely outperform specialized tools

None of this makes general LLMs weak extractors across the board. They win outright under a specific, identifiable set of conditions.

On semi-structured documents that follow a consistent template, paired with well-designed prompts, general-purpose LLMs can beat classical machine learning approaches on extraction accuracy without any task-specific fine-tuning at all. I The clearest evidence for this comes from a study by Gómez and Sánchez, who tested Gemini 1.5 Pro and Mistral-small on Spanish electricity invoices from the IDSEM dataset, running 19 parameter configurations against 6 different prompting strategies. The best setup, few-shot prompting combined with cross-validation, pushed both models to very high F1 scores. E Prompt quality, not hyperparameter tuning, drove the gap behind that headline number. N Zero-shot prompting and the best few-shot strategy differed by more than 19 percentage points, while changing the parameter configuration barely moved the needle. Document template consistency, not model architecture, was the biggest factor separating easy extractions from hard ones.

General LLMs also hold their own on scanned documents with degraded image quality, dense tables, and handwriting: LLM-based OCR beats traditional OCR engines by a meaningful margin on accuracy in these conditions. And for tasks that are fundamentally about meaning rather than position, summarizing a contract, identifying a clause, classifying a document's intent, a general LLM is the right tool for the job. A specialized extraction engine adds nothing extra there.

Most production document pipelines in 2026 run both kinds of systems side by side, routing documents to whichever one fits the task. J The split is the architectural argument, stated precisely. It's the argument, stated precisely: layout awareness matters enormously on some document types and barely at all on others, and the electricity invoice results hold specifically because invoices in that study kept a fairly consistent template structure, which is exactly the condition that falls apart in the failure modes already described: long documents, cross-page tables, and degraded scans.

What fine-tuning contributes, and leaves unsolved

A natural objection follows: if a general LLM gets fine-tuned on a company's own documents, doesn't that close the gap with a purpose-built, layout-aware system? Fine-tuning improves one thing substantially while leaving a separate problem untouched.

What it improves is schema adherence, the model's ability to output something that fits a fixed taxonomy correctly and consistently. A domain-tuned model, for instance, showed a large jump in correctly sorting documents into the right subcategory compared to a general-purpose model working from prompts alone. TorchSight's result here is a genuinely strong one and worth taking seriously on its own terms. B It measures correctly classifying which category a document belongs to, not resolving where a value sits on a page.

That distinction matters because fine-tuning doesn't change how the model reads a document. It still processes the page as a sequence of words; it has simply learned, through many examples, what fields to expect and how to label them once it finds them. That's why fine-tuning helps most on documents with stable, mostly-text layouts, and helps least on exactly the document types that caused trouble in the second section: dense tables, multi-column layouts, fields that reference each other across pages.

There's also a maintenance cost that compounds over time. The argument that "enough fine-tuning data eventually closes the gap" runs into a practical wall: edge cases keep accumulating, document templates change on their own schedule, and every time they do, the fine-tuned model needs retraining to catch up. F A layout-aware architecture handles those same variations structurally, by reading position rather than memorizing examples.

R Calibrated confidence scores and raw accuracy

A confidence score only earns its place in a production pipeline if it reliably tells an engineer where the model is likely wrong. L Accuracy and confidence calibration can move in opposite directions across different models. A high accuracy number alone says very little about whether the confidence signal attached to it can be trusted.

One audit pipeline evaluation made that gap concrete: the model with the single highest extraction accuracy, 0.77 Weighted Overall Accuracy, was not the model with the best-calibrated confidence scores. A different model led on every confidence metric in the comparison, while a third model posted a competitive 0.76 WOA and still lagged significantly on AUROC, the metric that measures how well confidence predicts actual error.

F The gap is also visible inside a single model, field by field. Amount fields tend to carry higher confidence because their formatting stays consistent from document to document; date fields tend to carry lower confidence because their position on the page varies more. An overall confidence average flattens that difference out and hides it from whoever is reading the dashboard.

The medical document research confirms the same pattern from another angle. Yu, Weile, and Courtot's 2026 study found that image distortions introduced by fax transmission cut extraction performance sharply, with GPT's accuracy dropping notably on distorted documents compared to clean ones, and the model's confidence scores did not reliably signal that the input had degraded at all.

Q For a production routing system, the one deciding which extracted results pass straight through and which get kicked to a human reviewer, a miscalibrated confidence score does more damage than having no score. It hands false assurance to exactly the cases that most need a second look. The evaluation standard built to catch this problem scores each field individually, by type, using Weighted Overall Accuracy alongside AUROC as a separate confidence-quality metric.

How the medical and healthcare context sharpens every constraint

Healthcare document processing is where layout blindness, miscalibrated confidence, and data-residency rules all collide at once, turning an architecture choice into a question of compliance and patient safety rather than a benchmark score.

Fax machines are still a routine part of how many healthcare institutions move documents between offices, and every fax transmission introduces its own distortion and artifacts that degrade legibility further. It's an active constraint shaping which model can be trusted on a given intake pipeline today, not a legacy curiosity. Yu, Weile, and Courtot's medRxiv preprint, published in January 2026, tested a range of LLM-based extraction tools against a large set of mock medical documents. GPT 4.1-mini came out as the most effective model, with an average F1 of 55.6. The best-performing local model was Google's Gemma3 running on image inputs with zero-shot prompting, at an average F1 of 41.3. The fax-introduced distortions had a real, significant effect on extraction accuracy, a result the authors call out specifically. More surprising: changing the prompting strategy barely moved those numbers, which runs against the common assumption that a better-written prompt can recover accuracy lost to a degraded scan.

Data residency adds a separate, non-negotiable constraint on top of all this. A cloud-hosted LLM, as a remote API-only model on a consumer or non-enterprise tier without a signed BAA, is not usable for privacy-sensitive medical document processing under HIPAA requirements. Enterprise-tier cloud offerings, Azure OpenAI, AWS Bedrock, Google Vertex AI among them, can be used once a signed BAA is in place, and that single requirement is the main reason zero-retention configurations and on-premises deployments exist in healthcare settings at all.

TorchSight's design is worth a brief mention again here, not for its classification accuracy, already covered above, but for the residency logic behind it: it was built to keep sensitive documents from ever leaving an organization's own infrastructure, addressing a real paradox where cloud-based data-loss-prevention tools require uploading the exact sensitive documents they're supposed to protect, and even a contractual guarantee from a provider still routes that document through infrastructure the organization doesn't control.

The regulatory floor keeps rising, too. The EU AI Act's high-risk requirements, originally set to take effect sooner, were deferred under Regulation (EU) 2026/1744 and are now set to become active in December 2027, layering new audit-trail and access-control obligations on top of the SOC 2, GDPR, and HIPAA rules that already govern how sensitive documents get processed. None of this changes the accuracy numbers from earlier sections. Q It changes whether a given pipeline is legal to run.

G What a production-grade evaluation methodology measures

The accuracy and F1 numbers vendors publish in marketing material rarely measure what actually breaks once a pipeline hits production volume. A credible evaluation has to score each field individually, by type, against labeled ground truth.

Weighted Overall Accuracy scores each field on a continuous scale from 0 to 1, with F1 run alongside it as a binary pass/fail threshold. Building the ground-truth set matters just as much as choosing the metric: it needs poor scans, unusual layouts, and documents that have caused errors before, deliberately included. A test set made only of clean, representative documents will make any model look better than it performs once it meets real intake volume.

Once a system is live, the metrics that matter are straight-through processing rate, field-level accuracy against a labeled sample, the top exception reasons showing up in the review queue, processing time against latency SLAs, and how each new model or version change shifts all of the above. A short weekly review of exception cases often produces faster accuracy improvements than months spent retuning a model.

Confidence calibration needs its own separate evaluation, run apart from the accuracy numbers. AUROC measured against ground-truth labels tells an engineer whether a model's confidence scores actually predict its own errors. That property, more than raw accuracy, is what makes human-in-the-loop review routing work at all.

That evaluation approach points toward a tiered fallback design that holds up under real volume. Check for an embedded text layer first, since digital-native PDFs often already have usable text and can skip OCR entirely. M Route clean documents to a fast model, which accepts the result once confidence clears a set threshold. Escalate anything below that threshold to a stronger, slower model. Send critical field failures to a human reviewer rather than letting them pass on a guess.

Sources

  1. Security Document Classification with a Fine-Tuned Local Large Language Model: Benchmark Data and an Open-Source System
  2. Benchmarking LLM-based Information Extraction Tools for ...
  3. Information Extraction from Electricity Invoices with General-Purpose Large Language Models
  4. OCR or Not? Rethinking Document Information Extraction in the MLLMs Era with Real-World Large-Scale Datasets
  5. LLM OCR vs Traditional OCR: When AI Wins (and When It Doesn't) — 2026 Benchmark
Filed underDocument AI

More in Document AI