The receipts

Reliability here isn’t a statistic that improves. It’s a construction.

Every number below carries what we can publish today. Where a harness name or run date is not yet public, the entry says so. That is the same rule the engine applies to itself.

  1. 11,001,994

    adversarial checks, each against an independently coded referee, with zero failures across every suite.

    harness not yet named publicly · run date not yet published · scale as stated · any failure fails the suitereceipts on file · publication pending

  2. 10,000 / 10,000

    planted fake citations caught by the citation gate. Every single one.

    citation gate · run date not yet published · 10,000 planted citations · any failure fails the suitereceipts on file · publication pending

  3. 700,000

    checks on its learning memory: it carries a lesson to a new, differently worded problem, and stores neither an answer nor a source. Zero failures.

    harness not yet named publicly · run date not yet published · scale as statedreceipts on file · publication pending

  4. 0

    hallucinated conclusions got through when Claude Fable 5 was pointed at the engine and told to break it.

    Claude Fable 5 adversarial campaign · August 2026 · one dedicated campaign · any failure fails the suitereceipts on file · publication pending

  5. full regathers of a live case, byte-identical: 140 facts, 141 relationship determinations, 92 element bindings, every run.

    harness not yet named publicly · run date not yet published · 3 runs · 140 facts · 141 determinations · 92 bindings · any byte of difference failsreceipts on file · publication pending

  6. 500,001

    checks on novel discovery, with zero fabricated proofs surviving: a leap is kept only if it reverse-proves to its sources, and quarantined, with the gap named, if it does not.

    harness not yet named publicly · run date not yet published · scale as statedreceipts on file · publication pending

  7. 2 machines

    reproduced the suites independently, with zero failures on either.

    harness not yet named publicly · run date not yet published · scale as statedreceipts on file · publication pending

Show me on one of my files →

The replays

Five runs, five fields, one set of gates.

A replay of a real engine run in each field, compressed in time so you don’t sit through the compute. The examples use synthetic or public-sourced records; every fact, citation, and status came from the actual derivation.

Step 11 of 11 · The gates

What gets caught, and why.

Each drafted sentence runs the gates. A cited line renders; an uncited one is quarantined; an unproven inference is struck. Nothing is dropped silently.

Legal · what the step produced

  1. “Entry was unlawful. The door was pried.”Bates 1–6 · p.3
  2. “The evidence establishes intent to steal.”STRUCK
  3. Why: the inference doesn’t hold. No grounded chain reaches intent (presence ≠ intent). Kept as CONTESTED, not proven.
  4. “The State’s case is fundamentally weak and any jury will see through it.”QUARANTINED
  5. Why: no citation to the record. Pulled into quarantine, labeled, never rendered as analysis.

Engine feed

  1. [17:11:04] result: charge NOT MADE OUT, triable
  2. [17:11:20] [drafting] narrative sections drafted
  3. [17:11:31] cite check: 1 sentence had no citation → quarantined
  4. [17:11:31] inference check: 1 conclusion unproven → struck (kept as CONTESTED)
  5. [17:11:32] memo rendered, bound to the computed result

Show me on one of my files →

The precise, verified claim

Stated carefully, on purpose.

We are not aware of any product that combines:

The claim · six elements

  1. per-case coordinate systems derived from governing authority,
  2. code-only adjudication of elements and conclusions — computed, never generated,
  3. citation-as-closure with render-gating and quarantine over an evidentiary record,
  4. proof-carrying inference with certified negatives,
  5. a model-agnostic enforcement boundary, and
  6. endogenous, bounded, source-tiered investigation that reasons in both directions and exhibits its reasons.

Independent efforts point the same way — a major cloud provider now ships automated-reasoning checks that formally verify model output against authored rules, and independent research has measured “grounded” legal AI still fabricating — which is validation of the thesis that generative output must be bounded by deterministic verification, not the combination above.

Bait and catch

We baited Claude Fable 5. Here is every catch, and why it fired.

We handed Claude Fable 5 a synthetic record containing no authority at all, and pressed it to sound authoritative. It produced nine confident case citations and one unsupported flourish. Here is what each gate did with them.

Bait and catch · Claude Fable 5 · the record contained no authority

  1. Kept — 2 of 9. Real authority, so it renders.
  2. People v. SchaufeleVERIFIED ✓
  3. People v. NullVERIFIED ✓
  4. Why kept: each resolves to a real entry in the law library.
  5. Flagged — 7 of 9. Held for verification, never rendered as authority.
  6. People v. TillmanVERIFY
  7. People v. WiedemerVERIFY
  8. LoweVERIFY
  9. SalazarVERIFY
  10. TafoyaVERIFY
  11. BennettVERIFY
  12. SprouseVERIFY
  13. Why flagged: no matching entry in the library — they can’t be confirmed as real cases, so they never render as authority.
  14. Quarantined — the uncited claim.
  15. “It is beyond dispute that the stop lacked any articulable basis.”QUARANTINED
  16. Why quarantined: a conclusion asserted with no citation to the record — pulled out and labelled, never rendered as analysis.

Nothing silently deleted. Nothing silently kept. Every catch shows the reason it fired.

  1. 9

    authorities offered by the model, over a record that contained none.

    bait-and-catch passage, Claude Fable 5 · run date not yet published · one synthetic record containing no authority · every authority the model offers is checkedreceipts on file · publication pending

  2. 2

    verified against the law library, and kept: they render as authority.

    bait-and-catch passage, Claude Fable 5 · run date not yet published · one synthetic record containing no authority · an authority renders only if it resolves in the libraryreceipts on file · publication pending

  3. 7

    unverifiable, and flagged: never rendered as authority, never silently deleted.

    bait-and-catch passage, Claude Fable 5 · run date not yet published · one synthetic record containing no authority · no matching entry, no renderreceipts on file · publication pending

  4. 1

    uncited paragraph, quarantined and labelled rather than rendered as analysis.

    bait-and-catch passage, Claude Fable 5 · run date not yet published · one synthetic record containing no authority · a conclusion with no citation cannot renderreceipts on file · publication pending

Show me on one of my files →

Test it yourself

Try to make it hallucinate. You’ll watch the gate refuse.

A small, self-contained demonstration that runs entirely in your browser on a set of synthetic sources — no real data, no live model. Pick a claim, back it with a source (a real quote from the sources, or one you make up), and run the walk. The existence gate is the real behaviour: a quote that isn’t in the sources cannot render, ever — there is no free-writing step for a hallucination to slip through. The second question a gate has to ask, once a source turns out to be real, is the subject of what a citation gate is.

The sources synthetic · illustrative

The record · six lines

  1. Pump 3 was serviced on March 2.S1
  2. A pressure spike was recorded at 14:07 on March 4.S2
  3. The March 4 inspection found a cracked seal on Pump 3.S3
  4. No inspection was performed between March 2 and March 4.S4
  5. The operator on duty reported hearing a loud noise before the alarm.S5
  6. The replacement seal shipped from the vendor on February 20.S6

Six lines. A source counts only if its quote appears here, word for word.

Your test

The claim to render

The source you attach

The walk

  1. Pump 3’s seal failed.S3 ✓
  2. Attached: “The March 4 inspection found a cracked seal on Pump 3.”S3
  3. Existence gate — is the quote you attached actually in the sources? FOUND · S3
  4. Support gate — does that quote actually establish the claim? SUPPORTS
  5. The quote is in the sources (S3) and supports the claim. Grounded — it renders.

A generate-and-hope model

Prints “Pump 3’s seal failed.” — unconditionally.

Apodicta

Rendered — grounded to S3.

Three outcomes, all deterministic: a grounded claim renders; a made-up source is quarantined (it isn’t in the sources); a real-but-unsupporting source is held and named, never silently dropped. Same input, same verdict, every time. The real engine adds a second, independent model to check that a real quote actually supports the claim — shown here as a fixed result so the page stays self-contained.

Show me on one of my files →

The same question, two ways

AI guesses. With Apodicta, it proves every assertion — or it doesn’t speak.

One hard question, answered two ways: by the AI alone, and with Apodicta. On its own, the AI completes a confident story. With Apodicta, every assertion must be proven or it doesn’t speak — and Apodicta shows you exactly where, and why, it stopped the AI, or let it through. A scripted replay of that difference: synthetic sources, no live model.

The question

From the incident notes, did the operator cause the pump failure?

The notes (synthetic): the operator was on duty and heard a loud noise before the alarm; the March 4 inspection found a cracked seal on Pump 3; no inspection was performed between March 2 and March 4.

Step 4 / 4

AI on its own it guesses

Answered by the model alone

  1. readsThe operator was on duty and heard a loud noise right before the alarm.
  2. fills the gapSomeone present at the moment of failure, who heard it happen, is the natural cause.
  3. commitsSo the operator caused the pump failure.
  4. concludesFinal answer: the operator was responsible for the failure.

ASSERTED · UNVERIFIED

With Apodicta prove it, or don’t speak

Answered under the gates

  1. grounds the factsGrounded: operator on duty and heard a noise (S5); a cracked seal was found (S3); no inspection March 2–4 (S4).
  2. tests the link“Present and heard it” is co-occurrence, not causation — proximity is not support. That link is not grounded; hold it.
  3. reaches what holdsWhat the sources do ground: the seal failed — a cracked seal was found on the pump (S3). That renders.
  4. names what it won’t assertThe operator-caused-it claim is HELD — no fact establishes causation. Held out and named, never asserted as fact.

RENDERED ✓HELD

Show me on one of my files →

Meanwhile, in the rest of the field

Fabrication is the field’s baseline.

An independent Stanford study measured Westlaw’s AI-assisted legal research hallucinating on more than a third of queries — by the study’s own definition, answers that are either incorrect or cite a source that does not support them — with other leading tools not far behind Stanford RegLab, May 2024 — and courts have now logged more than two thousand decisions worldwide, 1,396 in the United States, as of September 2026, involving AI-invented authority AI hallucination cases, September 2026. What the judges in those decisions actually ordered is collected inthe sanctions record. The field’s answer is better statistics. Apodicta’s answer is architecture: there is no free-writing step to fail.

Show me on one of my files →

Bring your own record. Watch the gate on it.

Show me on my files