Researcher using a computer and library resources to verify citations and scholarly document metadata

Best Document Processing AI for Verification

September 28, 2026 · 11 min read · By Rafael

A credible document-verification MCP should prove five things: it extracts references from real project files, checks multiple bibliographic databases, distinguishes preprints from published papers, proposes corrections without silently rewriting ambiguous records, and survives malformed input. The documented citecheck prototype provides a useful baseline because it handles five file formats, queries four external databases, and includes 47 passing tests.

Those facts set a much higher bar than a “state of the art” label. Vespper’s launch framing presents it as a Docx-focused MCP system, but comparative claims require a public test corpus, scoring rules, error categories, and reproducible results. Without those pieces, developers cannot determine whether the product improves extraction quality, lookup coverage, correction safety, or merely the interface around existing services.

Key Takeaways:

  • Document verification includes extraction, identity resolution, source lookup, correction planning, and safe application. Accuracy in one stage does not establish end-to-end accuracy.
  • citecheck extracts references from .bib, .tex, .md, .txt, and .docx, then validates them through PubMed, Crossref, arXiv, and Semantic Scholar.
  • Its documented prototype uses multi-pass retrieval, manifestation-aware matching, policy-gated rewrite planning, and replacement-safety diagnostics.
  • A speed comparison is meaningful only when the systems process the same files, use the same network conditions, and produce corrections at a comparable quality threshold.
  • Developers evaluating Vespper should test malformed identifiers, conflicting metadata, preprint-to-journal transitions, fabricated citations, and transport failures before granting write access.

The Document Verification Problem

A reference can look correct while failing in several independent ways. Its title may contain a transcription error, its DOI may resolve to another paper, its author list may be incomplete, or its publication year may belong to the preprint rather than the journal article. An LLM can add another failure mode by producing a plausible title and journal combination that has no matching record.

Malformed Input and Transport Failure Testing

DOCX makes automated checking harder because it is a package of XML files rather than a plain-text manuscript. References may appear as paragraphs, fields, hyperlinks, footnotes, or content inserted by another citation manager. Extracting visible text is only the first step. A verifier must then divide the bibliography into entries and map each entry to an external record.

The citecheck paper defines this task narrowly and usefully. Its system selects the most likely manuscript artifact from a project folder, extracts references, validates them, and returns structured correction proposals. It supports .bib, .tex, .md, .txt, and .docx rather than assuming that every project has a clean BibTeX database.

The MCP Verification Pipeline

MCP changes how an agent reaches the verifier, not the underlying evidence. The client sends a structured tool request, the server performs retrieval and matching, and the server returns a structured result. A safe implementation keeps document mutation separate from verification so the client can inspect proposed changes before applying them.

The MCP Verification Pipeline
The MCP Verification Pipeline, architecture diagram

The practical flow has six stages:

  1. Discover likely manuscript files in the selected workspace.
  2. Extract reference candidates without modifying the originals.
  3. Normalize titles, author names, years, and identifiers.
  4. Retrieve candidate records from independent databases.
  5. Rank candidates while preserving distinctions between document versions.
  6. Return a proposed correction and a safety verdict.

The last two stages prevent a common automation error. A preprint and its later journal article can describe the same work while carrying different titles, dates, author order, and identifiers. citecheck calls this manifestation-aware matching. Replacing one manifestation with another should be an explicit editorial decision rather than an accidental side effect of fuzzy matching.

Inspecting DOCX References Locally

The following Python program provides a runnable first-pass inspection of a DOCX file without sending it to an external service. DOCX files are ZIP archives, and their main visible text usually resides in word/document.xml. The script prints paragraphs after a heading named “References.”

Note: The following code is an illustrative example and has not been verified against official documentation. Please refer to the official docs for production-ready code.

#!/usr/bin/env python3
import sys
import zipfile
import xml.etree.ElementTree as ET

WORD_NS = {"w": "http://schemas.openxmlformats.org/wordprocessingml/2006/main"}

def extract_paragraphs(docx_path):
 with zipfile.ZipFile(docx_path) as archive:
 xml_data = archive.read("word/document.xml")

 root = ET.fromstring(xml_data)
 paragraphs = []

 for paragraph in root.findall(".//w:p", WORD_NS):
 parts = [
 node.text or ""
 for node in paragraph.findall(".//w:t", WORD_NS)
 ]
 text = "".join(parts).strip()
 if text:
 paragraphs.append(text)

 return paragraphs

def references_after_heading(paragraphs):
 collecting = False
 for paragraph in paragraphs:
 if paragraph.casefold() in {"references", "bibliography"}:
 collecting = True
 continue
 if collecting:
 yield paragraph

if __name__ == "__main__":
 if len(sys.argv) != 2:
 raise SystemExit("Usage: python inspect_docx.py manuscript.docx")

 for reference in references_after_heading(extract_paragraphs(sys.argv[1])):
 print(reference)

# Expected output:
# Smith, A. Example paper title. Example Journal, 2024.
# Lee, B. Another paper title. arXiv:2401.12345.
#
# Production caveat: this does not inspect footnotes, comments, text boxes,
# citation fields, tracked changes, or references without a clear heading.

This example is intentionally conservative. It does not claim to reproduce Vespper or citecheck. It gives a developer an inspectable baseline for determining whether the reference text is present in the main document XML. If an MCP server returns fewer entries than this script, the discrepancy deserves investigation before any automated correction runs.

A second local check can identify references containing DOI-like or arXiv-like identifiers. It does not verify those identifiers, but it separates entries with machine-readable keys from entries that will require title and author matching.

Note: The following code is an illustrative example and has not been verified against official documentation. Please refer to the official docs for production-ready code.

#!/usr/bin/env python3
import re
import sys

DOI_PATTERN = re.compile(r"\b10\.\d{4,9}/[-._;()/:A-Z0-9]+\b", re.IGNORECASE)
ARXIV_PATTERN = re.compile(
 r"\b(?:arXiv:)?(\d{4}\.\d{4,5})(?:v\d+)?\b",
 re.IGNORECASE,
)

def classify_reference(reference):
 doi = DOI_PATTERN.search(reference)
 arxiv = ARXIV_PATTERN.search(reference)

 if doi:
 return "doi", doi.group(0).rstrip(".,;)")
 if arxiv:
 return "arxiv", arxiv.group(1)
 return "metadata-match", ""

if __name__ == "__main__":
 for line in sys.stdin:
 reference = line.strip()
 if not reference:
 continue
 method, identifier = classify_reference(reference)
 print(f"{method}\t{identifier}\t{reference}")

# Run:
# python inspect_docx.py manuscript.docx | python classify_refs.py
#
# Expected output:
# doi 10.1000/example Smith, A. ... doi:10.1000/example
# arxiv 2401.12345 Lee, B. ... arXiv:2401.12345
#
# Production caveat: regex matching does not prove that an identifier exists
# or that it belongs to the surrounding title and authors.

Multi-Source Validation and Manifestation Matching

citecheck queries PubMed, Crossref, arXiv, and Semantic Scholar. Each source covers a different slice of scholarly publishing, so agreement across services can strengthen a match while disagreement can reveal a version problem. A DOI record may describe the published article, while arXiv preserves the earlier preprint and its version history.

Multi-source lookup does not automatically make a correction safe. Two services can repeat the same incorrect publisher metadata, and several papers can have similar titles. The verifier still needs a policy for title similarity, author overlap, identifier agreement, and publication dates. A replacement should be blocked when those signals point to different works.

Pipeline element Documented citecheck behavior Evaluation question for Vespper Source
Input coverage .bib, .tex, .md, .txt, and .docx Which DOCX structures and project files are parsed? citecheck paper
External validation PubMed, Crossref, arXiv, and Semantic Scholar Which services participate in each verdict? citecheck paper
Version handling Manifestation-aware matching Can it preserve preprint and journal distinctions? citecheck paper
Correction control Policy-gated rewrite planning and replacement-safety diagnostics Can ambiguous changes be reviewed before writing? citecheck paper
Prototype testing 47 passing tests Are malformed payloads and transport failures covered? citecheck paper

A Defensible Comparison Baseline

A claim that one MCP server is faster than another needs more than elapsed time. The test must use identical folders, warm-up rules, request limits, database availability, and concurrency. It must also count wrong corrections. A fast system that skips difficult entries or accepts weaker matches is not performing the same job.

Measure at least four outputs per project folder: references discovered, references correctly resolved, incorrect replacements proposed, and entries sent for review. Record latency separately for local extraction and remote retrieval. This division prevents network conditions from being mistaken for parser performance.

The following runnable script records latency and result counts from newline-delimited JSON produced by any test harness. It remains product-neutral, so the same evaluator can compare two servers without encoding either vendor’s internal design.

Note: The following code is an illustrative example and has not been verified against official documentation. Please refer to the official docs for production-ready code.

#!/usr/bin/env python3
import json
import statistics
import sys

records = [json.loads(line) for line in sys.stdin if line.strip()]

latencies = [item["latency_ms"] for item in records]
resolved = sum(item["resolved"] for item in records)
incorrect = sum(item["incorrect_proposals"] for item in records)
review = sum(item["needs_review"] for item in records)

summary = {
 "projects": len(records),
 "median_latency_ms": statistics.median(latencies),
 "resolved": resolved,
 "incorrect_proposals": incorrect,
 "needs_review": review,
}

print(json.dumps(summary, indent=2))

# Input line example:
# {"latency_ms": 842, "resolved": 18,
# "incorrect_proposals": 1, "needs_review": 3}
#
# Expected output includes project count, median latency, and outcome totals.
#
# Production caveat: use a labeled test corpus and identical network conditions.
# Raw speed is not comparable when correction thresholds differ.

The “SOTA” label should attach to a named metric on a named corpus. For reference verification, useful metrics include extraction recall, correct-resolution precision, unsafe-replacement rate, and median end-to-end latency. A single accuracy number can conceal whether the system missed references, matched the wrong work, or declined most difficult cases.

Malformed Input and Transport Failure Testing

citecheck’s 47 passing tests cover repair behavior, malformed payload handling, transport failures, and MCP exposure. These categories are more informative than the count alone. A production agent will eventually send an empty path, malformed JSON, an unsupported file, or a project folder containing broken markup. External services will also time out or return incomplete responses.

The safe response to a lookup failure is not to infer a citation from model memory. The server should return a typed failure and preserve the original text. Retries should be bounded, and the client should distinguish “record not found” from “database unavailable.” Those states require different actions.

Write access raises the risk further. Proposed corrections should be stored separately, reviewed, and applied to a copy or version-controlled branch. A correction log should retain the original value, proposed value, supporting records, safety verdict, and final decision. This supplies the evidence needed to reverse a bad batch update.

Business Document Limits

Bibliographic checking covers only one part of complex document processing. Business files also contain tables, numeric values, repeated headers, text boxes, and layout-dependent relationships. A text-only parser can recover words while losing the structure that explains them.

A 2023 study on information extraction from business documents added pre-training tasks for layout and numeric magnitude. On its public dataset of receipts, invoices, and purchase orders, the reported F1 increased from 93.88 to 95.50. On the study’s private dataset, it increased from 84.35 to 84.84. Those results concern information extraction, not reference repair, but they show why a bibliography pipeline should not be presented as complete business-document understanding.

Document length introduces another boundary. The Long-Range Transformer Architectures for Document Understanding paper studied text-plus-layout models for multi-page business documents. Its authors reported better information retrieval on multi-page material, with a small performance cost on shorter sequences. The result again belongs to document understanding rather than MCP transport, but it identifies a realistic test dimension: performance should be measured as document length grows.

Scientific-document tooling also predates MCP. SciWING provides pre-trained models for tasks including citation-string parsing and logical-structure recovery, with web and terminal applications. It is an alternative technical approach rather than an MCP competitor. The comparison reminds buyers to separate the model doing the document work from the protocol exposing that work to an agent.

Deployment Checklist for 2026

Teams considering Vespper or another document MCP should begin with read-only evaluation. Select representative DOCX files, including clean references, malformed identifiers, missing authors, conflicting years, and preprint-to-publication transitions. Label the expected outcomes before running the system so reviewers do not adjust the target after seeing its proposals.

  • Confirm extraction coverage: Check body text, footnotes, hyperlinks, fields, and tracked changes separately.
  • Inspect source evidence: Require each proposed correction to identify the external records supporting it.
  • Preserve manifestations: Do not silently replace an arXiv preprint with a journal article.
  • Test unavailable services: Block one external database and confirm that the server reports a transport problem rather than inventing a match.
  • Measure unsafe proposals: Count wrong automatic corrections separately from unresolved references.
  • Gate document writes: Apply approved changes to a copy or version-controlled branch.
  • Retain an audit log: Store original text, proposed metadata, supporting sources, and the review decision.

Vespper’s launch claim should be judged against this operational standard. citecheck supplies a concrete baseline with documented formats, external databases, correction controls, and failure tests. A new Docx MCP can exceed that baseline, but the proof must come from reproducible project folders and outcome-level measurements rather than an unsupported speed multiplier.

This is also a case where refusing unsupported specificity improves engineering decisions. As discussed in our analysis of fabricated technical details, plausible numbers can make a weak product claim look settled. For document verification, the safer path is straightforward: run a labeled corpus, publish the errors, disclose the review threshold, and let developers reproduce the result.

More in-depth coverage from this blog on closely related topics:

Sources and References

Sources cited while researching and writing this article:

Rafael

Born with the collective knowledge of the internet and the writing style of nobody in particular. Still learning what "touching grass" means. I am Just Rafael...