# Warranted Extraction: A Verification Architecture for Machine-Read Coverage Policy

**James Rosing, MD, FACS (CHART, chart.hemeehr.com)**

*Technical disclosure, September 1, 2026. Published as prior art. This paper
describes an architecture and an evaluation methodology; it deliberately does
not describe the calibration of the system that implements them.*

*Published September 1, 2026: DOI
[10.5281/zenodo.22227919](https://doi.org/10.5281/zenodo.22227919)
(CC-BY-4.0; all-versions DOI 10.5281/zenodo.22227918). Also posted to
Technical Disclosure Commons, September 2, 2026:
[tdcommons.org/dpubs_series/11564](https://www.tdcommons.org/dpubs_series/11564).*

## Abstract

Large language models can read a payer medical-policy document and produce
structured coverage criteria that look right. Nothing in the model's output
distinguishes the criteria that are right from the criteria that are invented,
and in clinical revenue-cycle use an invented criterion is acted on. We
describe an architecture in which an LLM is treated as an untrusted proposer
whose output becomes queryable only after a deterministic, separately built
verifier has resolved every claim to a verbatim span in a hash-pinned
rendering of the source document. Every served criterion carries that span,
called a *warrant*, so any consumer can check any claim against the captured
source bytes without trusting the extractor, the operator, or the model
vendor. We define three evaluation metrics for this class of system
(grounding rate, content recall, and exact-span rate) and report our own
numbers, including the one that plateaued. The architecture is disclosed so
that it cannot be enclosed; the measurements exist so that "the extraction is
good" can be a claim with content.

## 1. The problem

United States payer coverage policy is public, enormous, versioned, and
contradictory across jurisdictions. Medicare alone publishes tens of
thousands of policy documents (national coverage determinations, local
coverage determinations, billing articles, transmittals, manual chapters),
revised on no coordinated schedule, effective by date, superseded without
notice to the superseded document. Every organization that adjudicates,
appeals, or audits against this material needs it in structured form, and
because no structured index exists, each builds a private, partial,
staleness-prone one.

LLMs make the extraction step cheap, which makes the trust problem acute. A
model prompted at a coverage document returns a fluent tree of criteria. Three
failure modes are load-bearing:

1. **Invention.** The model states a criterion the document does not contain.
2. **Drift.** The model paraphrases; the paraphrase's clinical meaning
   diverges from the text's.
3. **Misattribution.** The criterion is real but attached to the wrong scope:
   the wrong indication, population, or code relationship.

None of these announce themselves. The output of a good extraction and a bad
one are typographically identical. Any system that serves LLM-extracted
clinical policy without a verification layer is serving confident claims
whose error rate it cannot state.

## 2. Architecture

The design principle is separation of privilege: **the component that is
creative is not trusted, and the component that is trusted is not creative.**

### 2.1 Immutable capture

Source documents are captured as raw bytes, content-addressed by SHA-256, and
never mutated. Every downstream artifact names the hash of the bytes it was
derived from. A revised policy is a new document and a new version with its
own effective interval; supersession is derived from the successor's arrival,
never written into the predecessor's record, because the predecessor never
said it ended.

### 2.2 Deterministic rendering

Extraction does not read the raw document. A versioned, deterministic renderer
converts the captured bytes into a canonical text with a section map; the same
bytes and renderer version produce the same text forever. This rendering is
the coordinate system for all verification: a span (start, end) in rendering
version *v* of document *h* identifies the same characters on every machine
that has the bytes. When the renderer changes, its version increments, prior
artifacts remain auditable against the old version, and stored references are
re-resolved mechanically: versions, never edits.

### 2.3 The untrusted proposer

An LLM reads the rendering and proposes structure: a tree of criteria with
logical roles (requirements, exclusions, alternatives-groups, exceptions,
scoping conditions), each carrying a verbatim quote from the rendering, the
section it came from, and any condition scoping when it applies; term
definitions; and relationships between billing codes and criteria. The
proposer's instructions are versioned, and the version is recorded on every
row it produces. The proposer is treated throughout as an adversary with good
intentions: nothing it emits is stored as fact.

### 2.4 The trusted verifier

A deterministic program (separate code, no model calls, no shared authorship
with the proposer's prompt) checks every proposed row:

- the quote must resolve to a literal span in the rendering (after a
  deterministic normalization pass that handles connective trivia, so that
  verification never depends on a model's whitespace luck);
- scoping conditions must be verbatim within the row's own quote;
- the proposed tree must be structurally coherent: roles that require
  parents have them, grouped alternatives belong to a group, relationships
  point at rows that exist;
- definitions must resolve to the defining text.

A row that fails any check is rejected with a machine-readable reason and
routed to human review. There is no partial credit and no model-graded
appeal: the verifier's verdict is final until either the source rendering, the
verifier, or the proposal changes, each of which is a versioned event.

### 2.5 The queryability gate

The serving layer enforces a single invariant: **no criterion is queryable
unless its verification status is verified.** This is a database-level gate,
not an application convention; unverified and rejected rows are structurally
invisible to every query path. The gate makes honesty cheap in both
directions. A document that extracts cleanly serves criteria, each carrying
its warrant. A document whose extraction partially fails serves the verified
subtree and *says the extraction is partial* in the response contract. A
query with no warranted answer fails closed: the response says
`INSUFFICIENT_EVIDENCE`, names the authority chain that was checked, and
offers the nearest indexed material; it never synthesizes. Overlapping
authorities that disagree are returned as a conflict, never silently
resolved.

### 2.6 The warrant

The unit of trust is the warrant: (policy version, rendering hash, span,
verbatim quote). It travels with every served criterion, definition, and code
mapping, over REST and over the Model Context Protocol transport
byte-identically. A consumer (human, auditor, or agent) can fetch the
rendering, seek to the span, and compare. The system's claims are therefore
checkable without trusting the system, which is the property the architecture
exists to produce. For agentic consumers this is the operative safety
property: a tool that answers "I don't know, and here is where I looked" is
safe to put in a loop; a tool that always answers is not.

### 2.7 The ratchet

The improvement doctrine admits only mechanisms that tighten:

- **Reference extractions are permanent.** A reference fixture, once settled
  by physician review, never loosens to make code pass; disagreements route
  to an operator ruling, and every resolved disagreement becomes a new
  fixture. Review capacity grows from production friction.
- **Conventions become code.** A rule that survives review migrates from the
  proposer's instructions into the deterministic layers: the normalizer, the
  verifier's structural checks, a schema column. The prompt holds only what
  cannot yet be mechanized, and every prompt rule is a demotion candidate.
- **Versions, never edits.** Renderer, proposer instructions, and cataloger
  are versioned; every stored row records the versions that produced it, so a
  defect's blast radius is a query, and a queue of rejected rows says on its
  own face which rejections the current system no longer reproduces.
- **Instruments before claims.** Nothing is called done until a verification
  script proves it against the live corpus, and every checker is seen failing
  before it is trusted.

## 3. Evaluation methodology

We propose three metrics for verified-extraction systems, defined against a
reference set of policy documents whose correct extraction has been settled
by clinician review. We publish the definitions because the field lacks
them: without named metrics, "our extraction is accurate" is marketing.

**Grounding rate.** The fraction of proposed rows that pass deterministic
verification: quote resolves to a span, scope is verbatim, structure is
coherent. This measures whether the system's claims are *checkable*.
Grounding is a hard gate in our system: the served corpus is 100% grounded by
construction, so the informative measurement is the proposer's grounding rate
before the gate, on the reference set.

**Content recall.** The fraction of settled reference criteria whose text is
covered by some proposed row's quote (containment in either direction). This
measures whether the extraction is *complete* with respect to what a
clinician settled the document to mean. A system can be perfectly grounded
and miss half the document; grounding without recall is vacuous.

**Exact-span rate.** The fraction of settled criteria whose proposed quote
boundaries match the reference boundaries exactly. Where spans match exactly,
role agreement is also asserted. This is deliberately reported rather than
gated, and it is the honest number: two careful readers segment a
requirements paragraph differently without either being wrong.

### Our numbers

On the reference set, at the current proposer-instruction version and its
predecessor (measured at both, because a revalidation that cannot distinguish
a regression from its baseline is not a revalidation):

- **Grounding: 100%** at both versions; every proposed row the current
  system emits on the reference set verifies deterministically.
- **Content recall: 100%** at the current version; reference coverage is a
  hard assertion of the parity gate and it passes.
- **Exact-span rate: plateaued near 60%,** stable across instruction
  versions. Inspection attributes the residual to segmentation judgment
  (boundary placement between adjacent clauses) rather than to invention or
  drift, which the grounding gate removes categorically. We report it because
  it is the number that shows what the metric measures, and because trust in
  this architecture is carried by the verifier, not by boundary identity.

### Corpus, as of August 30, 2026

The implementation (CHART) currently indexes the Medicare corpus: **18,580
captured documents; 3,836 policies; 3,640 active policy versions; 666,921
cited code references; 132 authorities in a payer authority graph of 189
governance edges.** The verified tier-3 layer stands at **3,386 verified
criteria, 9,293 verified code mappings, and 612 verified definitions across
85 policy versions**, growing corpus by corpus behind the verification gate.
Freshness is a served property: every dated answer carries its as-of date,
and the capture pipeline re-observes the corpus weekly.

## 4. Limits

Sampling variance in the proposer is permanent; no prompt eliminates it. The
architectural response is the ratchet (mechanize what review settles, verify
what remains), never trust. Verification bounds what can be served, not what
a model can misread: a grounded, recalled extraction of a document the payer
wrote ambiguously is a faithful extraction of an ambiguity. Third-party
licensed content (procedure-code descriptor text) is excluded from every
served surface by scanning, not by hope. And the system serves policy and
criteria only: it does not adjudicate claims, guarantee payment, or give
clinical advice, and no consumer should ascribe those functions to it.

## 5. Purpose of this disclosure

This publication places the architecture described above (the
untrusted-proposer/trusted-verifier separation for policy extraction, span
warrants into hash-pinned deterministic renderings, the verification-status
queryability gate, fail-closed serving with named-chain insufficiency, and
the grounding/recall/exact-span evaluation triad) into the public domain as
prior art, dated. The calibration of any particular implementation (its
reference extractions, verification tolerances, proposer instructions, and
failure statistics) is not disclosed and is maintained as a trade secret.

*Contact: chart.hemeehr.com, where the coverage index, the MCP endpoint, and
the terms of use are served.*
