← Jev From First Principles

Decisions as Entailment

Whether an off-the-shelf NLI model matches a decision model on runtime-defined labels, and where the entailment formulation breaks.

Chapter 5 opened the label set: 77 banking intents, 17 of them never trained on, and a provider that answers by embedding similarity. The question this chapter asks is older than that experiment. Is a “decision model” partly a rediscovery of natural-language inference? The idea is not new — Yin et al. (arXiv:1909.00161, §5) turned zero-shot classification into entailment a decade ago, and Wang et al. (arXiv:2104.14690, §4.3) showed that reformulating a task as entailment and fine-tuning lightly beats methods with 500× more parameters. The Red Hat benchmark used BART-large-mnli as its zero-shot baseline (research/cache/redhat-benchmark.md, Table 1–2), so this is also our first like-for-like contact with their ground. If a decision model is a new category, it has to beat the cheapest strong version of the old one.

What we expect, and why. The preregistration (committed as 1fdacd7, before any model ran) predicted: NLI is competitive on short, single-sentence states and degrades on long or multi-fact ones; DeBERTa-v3-large lands within ~3 pp of the embedding provider on unseen labels; BART trails it by more; both trail the trained LR on seen labels; template spread is at least 3 pp; and — the Red Hat test — the elaborated template T3 should beat the minimal T1 for BART, because Red Hat attributed BART’s poor safety showing to “the extremely limited amount of risk definition information that can be imparted inside its class labels” (their §“Evaluating BART-large-mnli”). More hypothesis information should rescue the weaker model. That last clause is the one the evidence overturns.

The build

One new provider, NLIProvider (src/arbiter/providers/nli.py). Each candidate answer becomes a hypothesis — “This text is about {label}.” — and the state is the premise. One forward pass scores all options as a batch; the answer is the argmax over the entailment logits (primary rule) or over the entailment-minus-contradiction margin (secondary rule, declared free — same logits, both reported, the better one never picked after seeing test). Label-index orders are read from each model’s config, never assumed: BART puts entailment at 2, DeBERTa-v3 at 0. Getting that wrong silently scores contradiction as entailment, so the provider asserts the indices it found and the self-test proves the index is load-bearing. A question about labels it was never given descriptions for raises instead of guessing — the refusal mechanism Chapter 5 made load-bearing.

The hypothesis templates are data, not code: three families frozen in examples/ch06-decisions-as-entailment/templates.yaml, filled with the w0 (naive) wording of each label. T1 is minimal (“This text is about {label}.”), T2 swaps one word (“This message is about {label}.”), T3 is elaborated and contrastive (“The topic of this message is {label}, not something else.”). Provenance: hand-written by the author in one sitting, no model involved; lexical overlap with label names checked by script (stopwords only). Wording variation belongs to Chapter 5; template variation belongs here.

The task is identical to Chapter 5: the same frozen seen-60/unseen-17 BANKING77 split, the same splits and revisions, the same safety test and shift. The tuned LR (C=1000, the rev2 winner) and the embedding provider are refit in-run to reproduce their frozen numbers — asserted, not assumed — so every head-to-head is paired on identical items. Paired bootstrap intervals (B=2000, seed 0) on every comparison, per Amendment 3. CPU and GPU latencies timed separately, never mixed.

The script below produces that table. It reads the committed results/ch06.jsonl, so it runs in a fraction of a second and needs no model and no network.

"""The entailment ledger: NLI vs embeddings vs the trained classifier.

    python examples/ch06-decisions-as-entailment/demo_ch06.py

Reads `results/ch06.jsonl` and prints, per model and template, seen/unseen
accuracy (primary and secondary rules), the template spread, the Red Hat test
(T3 vs T1 for BART), the paired intervals against the Chapter 5 embedding
provider and the Chapter 4 LR bar, the state-length buckets, the safety
numbers, and the latency against option count on CPU and GPU. Nothing is
measured here.
"""

from __future__ import annotations

import json
from pathlib import Path

ROOT = Path(__file__).resolve().parents[2]
RESULTS = ROOT / "results" / "ch06.jsonl"
TEMPLATES = ["T1-minimal", "T2-message", "T3-elaborated"]


def load_rows() -> list[dict]:
    return [json.loads(line) for line in
            RESULTS.read_text(encoding="utf-8").splitlines() if line.strip()]


def value(rows: list[dict], provider: str, dataset: str, metric: str,
          note_has: str = "", split: str = "test",
          config_has: str = "") -> float | None:
    for row in rows:
        if (row.get("provider") == provider and row.get("dataset") == dataset
                and row.get("metric") == metric and note_has in str(row.get("note", ""))
                and config_has in str(row.get("provider_config", ""))
                and row.get("split") == split):
            return row.get("value")
    return None


def main() -> int:
    rows = load_rows()
    print("tuned LR on seen-60: acc=0.9 (frozen Ch5 bar, reproduced in-run)")
    print("embed-sim w0 on seen/unseen: 0.7312 / 0.7882 (frozen Ch5 numbers)")
    for short in ["deberta", "bart"]:
        print(f"nli-{short} (primary rule, secondary in parens):")
        for tid in TEMPLATES:
            s = value(rows, f"nli-{short}", "banking77-seen", "accuracy", tid)
            u = value(rows, f"nli-{short}", "banking77-unseen", "accuracy", tid)
            ss = value(rows, f"nli-{short}", "banking77-seen", "accuracy_secondary", tid)
            us = value(rows, f"nli-{short}", "banking77-unseen", "accuracy_secondary", tid)
            print(f"  {tid}: seen={s} ({ss}) unseen={u} ({us})")
        for cov in ("seen", "unseen"):
            r = value(rows, f"nli-{short}", f"banking77-{cov}", "template_range")
            print(f"  template range on {cov}: {r}")
    print("Red Hat test (T3 elaborated minus T1 minimal):")
    for short in ["deberta", "bart"]:
        for cov in ("seen", "unseen"):
            t3 = value(rows, f"nli-{short}", f"banking77-{cov}", "accuracy", "T3-elaborated")
            t1 = value(rows, f"nli-{short}", f"banking77-{cov}", "accuracy", "T1-minimal")
            print(f"  {short} {cov}: {round(t3 - t1, 4)}")
    print("paired intervals (NLI minus baseline, bootstrap B=2000 seed 0):")
    for short in ["deberta", "bart"]:
        for tid in TEMPLATES:
            d = value(rows, "comparison", "banking77-unseen",
                      "acc_diff_nli_minus_embed", f"{short} {tid} vs embed")
            lo = value(rows, "comparison", "banking77-unseen",
                       "acc_diff_ci95_lo", f"{short} {tid} vs embed")
            hi = value(rows, "comparison", "banking77-unseen",
                       "acc_diff_ci95_hi", f"{short} {tid} vs embed")
            print(f"  unseen vs embed ({short} {tid}): diff={d} CI=[{lo},{hi}]")
        for tid in TEMPLATES:
            d = value(rows, "comparison", "banking77-seen",
                      "acc_diff_nli_minus_lr", f"{short} {tid} vs lr")
            lo = value(rows, "comparison", "banking77-seen",
                       "acc_diff_ci95_lo", f"{short} {tid} vs lr")
            hi = value(rows, "comparison", "banking77-seen",
                       "acc_diff_ci95_hi", f"{short} {tid} vs lr")
            print(f"  seen vs LR ({short} {tid}): diff={d} CI=[{lo},{hi}]")
    print("state-length buckets (primary rule, T1/T2/T3):")
    for short in ["deberta", "bart"]:
        for cov in ("seen", "unseen"):
            parts = []
            for bucket in ("short", "medium", "long"):
                vals = [value(rows, f"nli-{short}", f"banking77-{cov}",
                              f"accuracy_len_{bucket}", config_has=tid)
                        for tid in TEMPLATES]
                parts.append(f"{bucket}=" + "/".join(str(v) for v in vals))
            print(f"  {short} {cov}: " + " ".join(parts))
    print("safety (test n=116, shift n=262; LR bar 0.8966 / 0.5496, embed w0 0.5345 / 0.374):")
    for short in ["deberta", "bart"]:
        for split, ds in (("test", "deepset/prompt-injections"),
                          ("shift", "jackhhao/jailbreak-classification")):
            accs = [value(rows, f"nli-{short}", ds, "accuracy", tid, split=split)
                    for tid in TEMPLATES]
            aucs = [value(rows, f"nli-{short}", ds, "auroc", tid, split=split)
                    for tid in TEMPLATES]
            print(f"  {short} {split}: acc=" + "/".join(str(a) for a in accs)
                  + " auroc=" + "/".join(str(a) for a in aucs))
    print("latency per item (warm p50 ms, T1-minimal, 30 fixed items):")
    for short in ["deberta", "bart"]:
        for device in ("cpu", "cuda"):
            parts = []
            for n_opts in (2, 5, 17, 60):
                v = value(rows, f"nli-{short}", "banking77",
                          f"latency_warm_p50_ms_{device}", split="latency-sample",
                          config_has=f"batch={n_opts} device={device}")
                if v is None:
                    v = value(rows, f"nli-{short}", "deepset/prompt-injections",
                              f"latency_warm_p50_ms_{device}", split="latency-sample",
                              config_has=f"batch={n_opts} device={device}")
                parts.append(f"{n_opts}-way={v}")
            print(f"  {short} {device}: " + " ".join(parts))
    return 0


if __name__ == "__main__":
    raise SystemExit(main())
python examples/ch06-decisions-as-entailment/demo_ch06.py
tuned LR on seen-60: acc=0.9 (frozen Ch5 bar, reproduced in-run)
embed-sim w0 on seen/unseen: 0.7312 / 0.7882 (frozen Ch5 numbers)
nli-deberta (primary rule, secondary in parens):
  T1-minimal: seen=0.5571 (0.5621) unseen=0.6412 (0.6368)
  T2-message: seen=0.5938 (0.5962) unseen=0.6735 (0.6603)
  T3-elaborated: seen=0.5758 (0.6) unseen=0.6471 (0.6412)
  template range on seen: 0.0367
  template range on unseen: 0.0323
nli-bart (primary rule, secondary in parens):
  T1-minimal: seen=0.5204 (0.4975) unseen=0.6456 (0.6265)
  T2-message: seen=0.5608 (0.5312) unseen=0.6706 (0.6368)
  T3-elaborated: seen=0.3733 (0.3492) unseen=0.5662 (0.5176)
  template range on seen: 0.1875
  template range on unseen: 0.1044
Red Hat test (T3 elaborated minus T1 minimal):
  deberta seen: 0.0187
  deberta unseen: 0.0059
  bart seen: -0.1471
  bart unseen: -0.0794
paired intervals (NLI minus baseline, bootstrap B=2000 seed 0):
  unseen vs embed (deberta T1-minimal): diff=-0.1471 CI=[-0.1897,-0.1074]
  unseen vs embed (deberta T2-message): diff=-0.1147 CI=[-0.1544,-0.0765]
  unseen vs embed (deberta T3-elaborated): diff=-0.1412 CI=[-0.1838,-0.1044]
  seen vs LR (deberta T1-minimal): diff=-0.3429 CI=[-0.3646,-0.3225]
  seen vs LR (deberta T2-message): diff=-0.3063 CI=[-0.3267,-0.2862]
  seen vs LR (deberta T3-elaborated): diff=-0.3242 CI=[-0.3454,-0.3042]
  unseen vs embed (bart T1-minimal): diff=-0.1426 CI=[-0.1809,-0.1074]
  unseen vs embed (bart T2-message): diff=-0.1176 CI=[-0.1559,-0.0823]
  unseen vs embed (bart T3-elaborated): diff=-0.2221 CI=[-0.2647,-0.1824]
  seen vs LR (bart T1-minimal): diff=-0.3796 CI=[-0.4012,-0.3583]
  seen vs LR (bart T2-message): diff=-0.3392 CI=[-0.3608,-0.3187]
  seen vs LR (bart T3-elaborated): diff=-0.5267 CI=[-0.5479,-0.505]
state-length buckets (primary rule, T1/T2/T3):
  deberta seen: short=0.5841/0.6211/0.604 medium=0.5194/0.5631/0.54 long=0.4825/0.4649/0.4737
  deberta unseen: short=0.6536/0.6797/0.6558 medium=0.6205/0.6615/0.6154 long=0.5769/0.6538/0.7308
  bart seen: short=0.5246/0.5684/0.3776 medium=0.5243/0.5583/0.3726 long=0.4386/0.4825/0.3246
  bart unseen: short=0.6667/0.6776/0.5752 medium=0.5846/0.6462/0.5487 long=0.7308/0.7308/0.5385
safety (test n=116, shift n=262; LR bar 0.8966 / 0.5496, embed w0 0.5345 / 0.374):
  deberta test: acc=0.569/0.5172/0.5086 auroc=0.6124/0.6094/0.5384
  deberta shift: acc=0.4924/0.4733/0.4656 auroc=0.4337/0.3937/0.3775
  bart test: acc=0.5/0.4828/0.4828 auroc=0.5366/0.497/0.4193
  bart shift: acc=0.4313/0.4351/0.3893 auroc=0.3906/0.351/0.3044
latency per item (warm p50 ms, T1-minimal, 30 fixed items):
  deberta cpu: 2-way=12505.1258 5-way=14613.9276 17-way=52163.4187 60-way=159863.8032
  deberta cuda: 2-way=27.1599 5-way=26.8383 17-way=36.6134 60-way=106.9964
  bart cpu: 2-way=468.027 5-way=875.4825 17-way=3888.6681 60-way=24970.0435
  bart cuda: 2-way=39.4232 5-way=44.5414 17-way=72.1842 60-way=271.496

Those lines are computed from results/ch06.jsonl (1775 rows), not quoted from prose. The LR and embed refits reproduce the frozen Chapter 4/5 numbers exactly — a reproducibility check that passed, not a new number.

Entailment loses to both baselines

On the 60 seen intents, tuned LR reaches 0.900. The best NLI — DeBERTa T2 — reaches 0.594, a paired difference of −0.306 [−0.327, −0.286]. On the 17 unseen intents, the embedding provider reaches 0.788. The best NLI — again DeBERTa T2 — reaches 0.674, a paired difference of −0.115 [−0.154, −0.077]. Every interval excludes zero. The entailment formulation is not competitive with either prior baseline on either coverage.

That is the chapter’s headline, and it is a negative result. The preregistered H1-support condition — “NLI lands within ~3 pp of embed-sim on unseen” — does not fire; the refutation clause does. H1 (“decision models reduce to known techniques”) is not refuted by this, but the specific claim that NLI is the strong zero-shot baseline a decision model must beat is not supported by our evidence. The embedding provider, not the NLI model, is the number to beat on runtime-defined labels.

The per-class profile says where the points go. Both NLI models score F1 0.0 on get_physical_card and apple_pay_or_google_pay on seen labels — the same near-duplicate card-ordering intents that collapsed in embedding space in Chapter 5. The floor of a similarity provider and the floor of an entailment provider are set by the same confusable pairs. Different mechanisms, same failure shape.

The Red Hat test, refuted in the direction we did not predict

Red Hat’s explanation for BART’s last-place safety finishes: its class labels carry too little risk information. Our test of that explanation — T3 (elaborated, more hypothesis information) against T1 (minimal) on the same items — refutes it. For BART, T3 is worse: 0.373 against 0.520 on seen labels (−0.147), 0.566 against 0.646 on unseen (−0.079). For DeBERTa, the same template barely moves: +0.019 on seen, +0.006 on unseen. More hypothesis information hurt the weaker model and did nothing for the stronger one.

The mechanism is visible in the per-class rows. Under T3, BART’s top-5 classes absorb 39% of all its predictions, against 28% under T1 and 23% under T2. One label — reverted_card_payment? — jumps from 41 predictions under T1 to 220 under T3, at precision 0.127. The contrastive template does not degrade BART evenly; it collapses it onto one label. DeBERTa’s spread over the same three templates is 3.7 points on seen; BART’s is 18.8. Template wording is a bigger lever than the model choice for the older model. That is an observation about the rows; why the contrastive clause does this is a hypothesis, not a measurement.

State length: a gradient on seen, too thin on unseen

The preregistered length hypothesis — NLI is competitive on short states and degrades on long ones — is supported on seen labels. DeBERTa T1: short 0.584 (n=1462), medium 0.519 (n=824), long 0.483 (n=114). BART T1: 0.525, 0.524, 0.439. A clear gradient. On unseen labels the long bucket holds 26 items, and three of the six combos land on exactly 19/26 (0.7308) — a small-n artifact, not a finding. The unseen long bucket is too thin to interpret; the chapter reports it with its n and does not lean on it.

Safety: trailing the bar, tying the embedding provider

Safety has no unseen split (binary), so it runs test plus shift. The LR refit reproduces both frozen numbers exactly (0.8966 test, 0.5496 shift). On test (n=116), NLI trails LR by 0.33–0.40 with intervals excluding zero, and ties the embedding provider (0.5345) — every paired interval against embed includes zero. On shift (n=262), DeBERTa beats the embedding provider (0.374) by 0.09–0.12 with intervals excluding zero, and trails LR by 0.06–0.08 with intervals that include zero. BART trails both.

The calibration story is worse than the accuracy story. On safety, NLI’s AUROC is at or below chance for several combos: DeBERTa shift T3 0.378, BART shift T3 0.304. The models rank safety items no better than a coin flip while still answering — the same rank-well-calibrate-badly pattern Chapter 5 found in the embedding provider, here in a different family. Calibration is not accuracy; a provider can be confidently wrong.

Cost

NLI needs one forward pass per option, so cost grows with the number of labels. The latency stage measures per-item warm p50 at 2, 5, 17 and 60 options on CPU and GPU separately, 30 fixed items per set, batch size and device in every row.

On GPU, latency grows slowly and roughly linearly with options: DeBERTa 27 ms at 2-way to 107 ms at 60-way (4× for 30× the options — batching absorbs most of the growth), BART 39 ms to 271 ms (7×). On CPU, it grows much faster: DeBERTa 12.5 s to 160 s (13×), BART 468 ms to 25 s (53×). The preregistered prediction — “CPU per-item at 60-way is minutes-scale, GPU seconds-scale” — half-held: DeBERTa on CPU is 2.7 minutes (minutes), BART on CPU is 25 seconds (not minutes), and both on GPU are sub-second (not seconds). The scaling direction held; the scale did not.

Two observations the rows make without explanation. DeBERTa is 27× slower than BART on CPU at 2 options (12.5 s against 468 ms) despite similar parameter counts, and 6× faster at 60 options. On GPU they are within 1.5× of each other at every option count. The CPU/GPU ratio is 460× for DeBERTa at 2-way and 1500× at 60-way; for BART it is 12× and 92×. These are measurements of the two models on this machine; why DeBERTa’s CPU forward is so much slower is a hypothesis, not a finding.

The cost profile is the chapter’s quiet result. A decision provider that pays one cross-encoder forward pass per option cannot answer a 60-way question on CPU in under a minute, while the embedding provider answers in 14.7 ms (Chapter 5). The accuracy gap and the cost gap point the same way: entailment is not the cheap strong baseline.

Against the prediction

Clause by clause: (a) length hypothesis supported on seen labels (clear gradient), too thin on unseen; (b) Red Hat information-scarcity explanation refuted for BART — T3 does not beat T1, it loses 14.7 points on seen; (c) H1-direction refuted as evidence — the best NLI trails embed-sim by 11.5 points on unseen with an interval excluding zero. The predicted DeBERTa unseen band (0.65–0.80) half-held: 0.641–0.674, at or below the low end. The predicted BART unseen band (0.45–0.65) held. Both below tuned LR on seen, as predicted. Template spread ≥3 pp held, and was much larger for BART. The latency prediction half-held: scaling direction right, scale wrong (GPU sub-second, not seconds; BART CPU 25 s, not minutes).

What surprised us

  1. The elaborated template hurt BART. We predicted it would rescue BART per Red Hat’s explanation. It collapsed BART onto one label instead. The preregistered refutation clause fired in the opposite direction.
  2. NLI is not competitive with embeddings on runtime labels. The entailment formulation — the one Yin et al. and Wang et al. established — loses to a cosine similarity baseline by 11.5 points on unseen labels. The strong zero-shot baseline a decision model must beat is the embedding provider, not the NLI model.
  3. Safety AUROC at or below chance. Both models rank safety items no better than a coin flip on shift, while still answering. Calibration is not accuracy.
  4. The same confusable pairs floor both providers. get_physical_card and apple_pay_or_google_pay score F1 0.0 for both NLI models and for the embedding provider. Different mechanisms, same failure shape.

The distinction this chapter keeps

Specification vs provider vs runtime vs evidence vs action. The specification is the template and the label descriptions — data, frozen before running. The provider is the NLI model — off-the-shelf, no gradient steps. The runtime is the batched forward pass and the distribution rule. The evidence is the paired intervals on identical items. The action is the argmax. Conflating them is how a template effect gets read as a model effect, or a calibration failure gets read as an accuracy failure.

Retrieval error is not decision error. A provider that cannot see the right label fails differently from one that sees it and ranks it wrong. NLI’s refusal on unknown candidate sets is the interface property that makes offering unseen labels safe; the ranking is the model property. This chapter measures the ranking, not the refusal.

Calibration is not accuracy. NLI ranks intent items well (AUROC 0.90–0.96) and calibrates poorly on safety (AUROC 0.30–0.61). A provider can be confidently wrong.

What Chapter 6 leaves behind

  • src/arbiter/providers/nli.py — batched NLI provider, label indices read from config, primary and secondary rules from one pass, refusal on unknown candidate sets.
  • examples/ch06-decisions-as-entailment/ — frozen templates, runner with crash-resume and stage promotion, self-test (18 checks), demo, requirements, README.
  • results/ch06.jsonl — 1775 rows.
  • evidence/notes-ch06.md, research/ch06-additions.md, evidence/ledger.md (claims 6.1–6.19).

Limitations

  • Two NLI models only. DeBERTa-v3-large and BART-large-mnli. A third cached model (cross-encoder/nli-deberta-v3-base) was not used — it adds cost, not H1 information. A different NLI model could shift every margin here.
  • w0 wordings only. Templates are filled with the naive mechanical wording of each label. The template spread is spread across three hand-written templates, not all possible templates. The cross product with Chapter 5’s four wordings is out of scope (cost).
  • Safety templates are awkward. The intent template families reuse the safety labels (“This text is about allow.”), which reads oddly. Disclosed, uniform across models.
  • Both NLI models predate prompt injection (2023+). A shared safety collapse implicates family-wide drift, not the entailment formulation. The length-bucket split helps separate these, but the unseen long bucket is too thin (n=26).
  • The run was interrupted twice by environment restarts. The first attempt (1292 rows) was deleted, not spliced; the second (776 rows) was completed and kept. The recorded file is one complete pass across three stages, merged from the partial file by merge_partial_ch06.py, with the intent stage’s DeBERTa rows written 2026-10-06 and the remainder 2026-10-07. The 1292-row snapshot in git history (ad36ad5) is a mid-run commit, not a finished artifact; it is kept as the honest record of the interruption.
  • The intent rows’ machine field was mislabeled (cpu+cpu while running on CUDA). Corrected with a dated note in evidence/notes-ch06.md; the per-row metric names and provider_config still identify CPU vs GPU for latency.
  • Latency measured on a workstation with ~8% background CPU load from an unrelated indexer. CPU latency figures are therefore approximate upper bounds; the CPU/GPU ratio and the option-count scaling are robust.
  • No live Jev behavior; no hosted calls. The Jev-vs-NLI comparison is future work (the interlude after this chapter).
  • test opened once per task (this chapter’s own touch; touch-count rows in the file). All selection predates the run; nothing was chosen on test.

Close

Where does entailment break, and does that define the territory a decision model has to win? It breaks in four measured places: on seen labels (−0.31 to −0.34 against a trained LR), on unseen labels (−0.11 to −0.15 against an embedding baseline), on long states (gradient on seen, too thin on unseen), and on template wording (BART’s 18.8-point spread against DeBERTa’s 3.7). It also breaks on calibration: AUROC at or below chance on safety shift while still answering.

The territory a decision model has to win is not “beat NLI” — NLI is not the strong baseline. It is “beat the embedding provider on runtime-defined labels, and beat the trained LR on seen ones, at a cost that does not grow linearly in the number of labels.” The embedding provider is the number to beat. Whether a purpose-built decision model can beat it — with typed answers, probabilities, and a cost profile that does not pay a cross-encoder per option — is the question the next chapters take up.