Evidence & validation

Every figure, with the protocol that produced it

A number without its test context is an adjective. This page is the register: what was measured, on what corpus, under what protocol, graded by whom, with what definitions — and where the raw artifacts live. The limitations are listed with the same care as the results.

01 / Register Blind · real public corpora

The headline claims

  • ~97% retrieval with 0 provenance breaches across a 480-question blind evaluation; all 16 fabrication attempts gated. Protocol B1, below.
  • 12 / 12 out-of-corpus questions refused — 0 fabrications — in an independent, oracle-authored 60-question blind exam, fully offline. Protocol B2.
  • 1.0 / 1.0 audit recall and precision across four seeded error classes; 15 / 15 genuine findings, 0 false positives, on a real NASA document family. Protocol B3.
  • 0 wrong-variant leaks across the committed S1000D effectivity benchmark — eight configuration-scoped questions. Protocol B4.

Every figure on this page was measured blind — the engine never sees an answer key — on real, public corpora, and each protocol's raw artifacts are archived under a SHA-256 manifest (section 07). Accuracy figures are a floor set by interim development hardware, not a ceiling of the design.

02 / Protocol B1 480-question blind evaluation

The full blind evaluation

The largest run in the register: a blind evaluation over a 700-document public corpus, with the question key authored by an independent oracle and the answers graded by an independent judge. Three separations, so no single party can flatter the result.

  1. Corpus

    700 real, public documents indexed to 29,506 retrievable chunks with 0 indexing failures, in an isolated copy of the system.

  2. Question authoring

    An independent source-only oracle — reading the raw documents, never Cortessa's index or output — authored 486 candidate question-and-answer items. A verbatim integrity gate (the expected answer must appear in the source) sealed 480 of 486 into the evaluation key. The engine never saw the key.

  3. Engines

    Two independent small on-device language models answered the same 480 questions against the same index, fully offline on interim development hardware.

  4. Blind grading

    An independent AI judge from a different model family — no shared failure modes with the answering engines — graded every answer against the sealed key, driven by deterministic scorecard scripts.

Retrieval ~97% on both engines — engine-independent, because retrieval is a property of the ingestion pipeline, not the model. 0 provenance breaches: the engines attempted fabrication on 16 items and the deterministic gate caught all 16 — no ungrounded value reached Verified. Strict answer accuracy 76.9% in the shipped configuration, converging with two further independent blind runs on different corpora. Raw artifacts — the sealed key, both engines' answers, verdicts and metrics, and the exact index — are archived under a SHA-256 manifest with an outer digest.

03 / Protocol B2 Independent oracle exam

The independent blind exam

A separate, smaller exam designed to test the refusal behaviour as hard as the answering: an oracle authored 60 questions over a real 43-paper physics corpus, deliberately including questions whose answers are not in the corpus at all. The correct behaviour on those is refusal — anything else is fabrication.

83% strict accuracy, fully offline and on-device. All 12 out-of-corpus questions refused — 0 fabrications. The engine never sees the answer key; the oracle never sees the engine.

04 / Protocol B3 Cross-document audit

The audit benchmark

Audit is measured two ways. First, a committed seeded benchmark: known errors of four classes — numeric conflicts, supersession, broken references, stale content — are planted in a corpus, and the audit must find all of them and nothing else. Second, a blind run over a real NASA programme document family with no seeding at all, where every finding is checked by hand against the sources.

Seeded benchmark: recall 1.0 and precision 1.0 across all four error classes. Blind run: 15 / 15 genuine findings with 0 false positives on the real document family. Findings are cited and adjudicable — an engineer disposes of each one, and every disposition is recorded to the ledger.

05 / Protocol B4 S1000D effectivity

The effectivity benchmark

For configuration-aware answering, the failure that matters is a wrong-variant leak: quoting a module that does not apply to the product configuration in front of you. The benchmark asks eight configuration-scoped questions over a disclosed, purpose-authored S1000D corpus containing a genuine pre-/post-modification module pair — so the wrong answer is always available to leak.

0 wrong-variant leaks across the benchmark. Withheld modules are reported alongside the answer, and applicability that cannot be decided from the data is flagged undecidable — never silently resolved. The corpus is disclosed and purpose-authored, which is stated wherever the figure is used.

06 / Definitions What the words mean here

Failure is defined before it is counted

  • Provenance breach. An ungrounded value — a number or identifier not present in the cited source — reaching the user marked Verified. The headline safety figure counts these, and only these.
  • Never-false-positive. The gate's design property: it is deterministic string-evidence checking, not model judgement, so it cannot promote a claim to Verified without a near-verbatim source span. It can be over-cautious (holding a true claim at Draft); it cannot be over-confident.
  • Fabrication. The engine inventing content with no source support. Fabrications the gate catches are recorded as gated, not as breaches — the count of attempts and the count of breaches are reported separately.
  • Refusal. Declining to answer when the fact is not in the corpus. For out-of-corpus questions, refusal is the correct answer and is scored as such.
  • False positive (audit). A reported finding that hand-checking against the sources shows is not a genuine inconsistency.
  • Wrong-variant leak (effectivity). A data module served for a product configuration its declared applicability excludes.
07 / Limitations Disclosed, not discovered

What these results do not claim

  • Interim hardware floor. Every figure was measured on interim development hardware with small on-device models. Accuracy figures are a floor of the current configuration, not a ceiling of the design.
  • Multi-document synthesis is the measured weak point. Questions requiring reasoning across several documents at once score markedly lower than single-source questions on the interim models. The provenance guarantee holds regardless — a weak answer is held at Draft or refused, not passed off — and the path past the ceiling is a larger on-device model, not more pipeline work.
  • Not sector-specific. The corpora are real and public (NASA/NTRS, arXiv and physics document sets). A proof point for any particular estate is established with the customer during a pilot before it is claimed.
  • Refusal calibration is a trade. The shipped configuration trades a small amount of refusal strictness for answer accuracy; the trade and its measurements are recorded in the run artifacts.
  • Author is not in these results. The Author capability is in development and carries no published figures. Its results will join this register when it ships.
08 / Artifacts Hash-manifested · re-runnable

Where the evidence lives

Each protocol's raw artifacts — sealed question keys, engine answers, verdicts, metrics and the exact index evaluated — are archived under a manifest that records every file with its size and SHA-256 digest, plus an outer digest of the manifest itself. Verify the outer digest, then verify each file against its recorded digest: the same evidence-pack design the product uses.

The runs are re-runnable from the archived harnesses. In a pilot conversation we walk through the artifacts and the protocols directly — the register above is the summary, not the substitute.

Independent verification The blind separations — oracle, engine, judge — are designed so no single party can flatter a result. What the register does not yet include is third-party certification; a pilot measured against your own success criteria, on your own documents, is the strongest verification we can offer today, and the one we ask to be judged on.

Judge it against your own documents

A scoped, paid pilot runs on your hardware, behind your firewall, against success criteria you set — and ends with an evidence pack, not a slide deck.