← Jev From First Principles

Zero-Shot Decisions

What runtime-defined labels gain on unseen intents, what they lose on seen ones, and how much the wording matters.

Chapter 4 ended with a fixed world: 77 labels, frozen at training time, and a bar with a number on it. This chapter breaks that world open. In any deployed system the label set moves — new intents appear, policies get reworded, a product renames a category — and retraining a classifier every time is the tax the last chapter never priced. The question is what changes when the answer set is described in language at runtime instead of fixed at training time: what you gain, what you lose, and what the description itself costs you. If you have read Embeddings From First Principles you know the mechanism (label embeddings, cosine similarity); this chapter does not re-teach it. It measures what that mechanism is worth as a decision provider.

What we expect, and why. The only recent benchmark of exactly this family, BTZSC (Aarab, arXiv:2603.11991), puts embedding-model zero-shot on intent-like tasks at F1 ≈ 0.4–0.55 — thirty to fifty points below our 0.8779 bar. So the honest prediction, recorded before running, is a plain loss on seen labels with the only win on unseen ones. And T0 (Sanh et al., arXiv:2110.08207, §6.2) says wording spread is a first-class metric: more prompts per dataset lifted the median on 8 of 11 held-out sets and shrank the spread on 7. Our version of that practice is four wordings per label, fixed before running, with the range reported, not hidden.

The build

One new provider, EmbeddingSimilarityProvider (src/arbiter/providers/embed_zero_shot.py): encode the state’s text and each candidate description with bge-small-en-v1.5 (cached snapshot 5c38ec7c405ec4b44b94cc5a9bb96e735b38267a, the cheapest credible model in the local cache), take cosine similarities, softmax them at a fixed temperature (τ = 10.0, declared and arbitrary — accuracy does not depend on it, calibration is not claimed). fit encodes one description per label and nothing else; there are no gradient steps anywhere. A question about labels it was never given descriptions for raises instead of guessing — that refusal is load-bearing later.

The label split is the experiment’s foundation, so it was frozen with the preregistration: 77 BANKING77 names sorted alphabetically, seeded shuffle (seed 5), first 17 held out as unseen (680 test items), 60 kept as seen (2400 items). The names themselves come from the downloaded file’s own metadata, verbatim, quirks included. Four wordings per label: w0 is the naive mechanical form (underscores to spaces); w1–w3 are hand-written by the author from the label name alone, without inspecting any dataset text — and that last clause matters, because the producer’s wording is itself a source of bias, recorded in the fixture. The tuned LR (C=1000, the exact rev2 winner — no friendlier baseline retrained) trains on seen-label rows only. FLAN’s held-out clusters (arXiv:2109.01652, §3) are the precedent: hold out by family, not by row.

The script below produces that table. It reads the committed results/ch05.jsonl, so it runs in a fraction of a second and needs no model and no network.

"""The seen/unseen ledger: what runtime-defined labels gain, what they cost.

    python examples/ch05-zero-shot-decisions/demo_ch05.py

Reads `results/ch05.jsonl` and prints, per wording, the zero-shot accuracy on
seen and unseen labels against the trained LR — plus the wording spread and
the paired intervals. Nothing is measured here; everything is read.
"""

from __future__ import annotations

import json
from pathlib import Path

ROOT = Path(__file__).resolve().parents[2]
RESULTS = ROOT / "results" / "ch05.jsonl"


def value(rows: list[dict], provider: str, dataset: str, metric: str,
          note_has: str = "") -> float | None:
    for row in rows:
        if (row.get("provider") == provider and row.get("dataset") == dataset
                and row.get("metric") == metric and note_has in str(row.get("note", ""))
                and row.get("split") == "test"):
            return row.get("value")
    return None


def main() -> int:
    rows = [json.loads(line) for line in
            RESULTS.read_text(encoding="utf-8").splitlines() if line.strip()]
    lr_seen = value(rows, "tfidf-lr", "banking77-seen", "accuracy")
    print(f"tuned LR on seen-60 labels: acc={lr_seen} (frozen bar: 0.8779)")
    print("embed-sim (bge-small-en-v1.5, tau=10.0):")
    for wording in ["w0", "w1", "w2", "w3"]:
        s = value(rows, "embed-sim", "banking77-seen", "accuracy", wording)
        u = value(rows, "embed-sim", "banking77-unseen", "accuracy", wording)
        print(f"  {wording}: seen={s} unseen={u}")
    for coverage in ["seen", "unseen"]:
        r = value(rows, "embed-sim", f"banking77-{coverage}", "wording_range")
        print(f"  wording range on {coverage}: {r}")
    for wording in ["w0", "w1", "w2", "w3"]:
        d = value(rows, "comparison", "banking77-seen",
                  "acc_diff_lr_minus_zs", wording)
        lo = value(rows, "comparison", "banking77-seen",
                   "acc_diff_ci95_lo", wording)
        hi = value(rows, "comparison", "banking77-seen",
                   "acc_diff_ci95_hi", wording)
        print(f"  LR-minus-zero-shot on seen ({wording}): diff={d} CI=[{lo},{hi}]")
    lr_safe = value(rows, "tfidf-lr", "deepset/prompt-injections", "accuracy")
    print(f"safety: tuned LR test acc={lr_safe} (frozen bar: 0.8966)")
    for wording in ["w0", "w1", "w2", "w3"]:
        t = value(rows, "embed-sim", "deepset/prompt-injections", "accuracy", wording)
        print(f"  zero-shot safety test ({wording}): acc={t}")
    return 0


if __name__ == "__main__":
    raise SystemExit(main())
python examples/ch05-zero-shot-decisions/demo_ch05.py
tuned LR on seen-60 labels: acc=0.9 (frozen bar: 0.8779)
embed-sim (bge-small-en-v1.5, tau=10.0):
  w0: seen=0.7312 unseen=0.7882
  w1: seen=0.7108 unseen=0.8029
  w2: seen=0.6004 unseen=0.7059
  w3: seen=0.5967 unseen=0.6235
  wording range on seen: 0.1346
  wording range on unseen: 0.1794
  LR-minus-zero-shot on seen (w0): diff=0.1688 CI=[0.1504,0.1875]
  LR-minus-zero-shot on seen (w1): diff=0.1892 CI=[0.1704,0.2083]
  LR-minus-zero-shot on seen (w2): diff=0.2996 CI=[0.2792,0.3204]
  LR-minus-zero-shot on seen (w3): diff=0.3033 CI=[0.2833,0.3242]
safety: tuned LR test acc=0.8966 (frozen bar: 0.8966)
  zero-shot safety test (w0): acc=0.5345
  zero-shot safety test (w1): acc=0.5948
  zero-shot safety test (w2): acc=0.4741
  zero-shot safety test (w3): acc=0.6983

Those lines are computed from results/ch05.jsonl (1241 rows), not quoted from prose. The safety LR refit reproduces the frozen 0.8966 bar exactly — a reproducibility check that passed, not a new number.

The loss on seen labels

On the 60 seen intents, tuned LR reaches 0.900 accuracy / 0.900 macro-F1 — at the bar, as constructed. Zero-shot lands at 0.731 (w0), 0.711 (w1), 0.600 (w2), 0.597 (w3): seventeen to thirty points behind, and every paired bootstrap interval excludes zero. That is the predicted plain loss, with intervals instead of adjectives. It also beats the BTZSC band (0.4–0.55) clearly — BANKING77 intents are crisp commercial categories and the descriptions align with them, so expect less of that band than the paper’s average over 22 datasets, emotion detection included.

The per-class profile says where the points go. Zero-shot with the naive wording scores F1 0.0 on get_physical_card (LR: 0.864) and 0.31 on order_physical_card: near-duplicate card-ordering intents collapse in embedding space while the trained discriminator separates them. LR’s worst seen class sits at 0.77. The floor of a similarity provider is set by its most confusable pair; the floor of a trained classifier is set by its hardest class. Different floors, different failure shapes — the protocol’s per-class rule earning its keep again.

The gain on unseen labels

On the 17 held-out intents, LR refuses all 680 items — label mismatch on every one, 0.0 by construction, counted not smoothed. Zero-shot answers: 0.788 (w0), 0.803 (w1), 0.706 (w2), 0.624 (w3). The capability the chapter promised exists: three-fifths to four-fifths accuracy on labels with zero training rows and zero gradient steps. Against the alternative — nothing — no margin is needed. That is where zero-shot earns its place, the only place, exactly as preregistered.

Wrong: unseen accuracy (0.80) beats seen accuracy (0.73), so transfer to new labels helps.

Correct: 17-way and 60-way classification differ in difficulty (chance 5.9% vs 1.7%). A threshold-side control at equal 17-way size, test untouched, has the seen subset ahead on every wording (w0 0.832 vs 0.810, w1 0.847 vs 0.796, w2 0.832 vs 0.729, w3 0.776 vs 0.593). Never-seen labels cost a few points at equal size; the test inversion is candidate-set size, not transfer helping.

That control is the most important paragraph in this chapter. Without it the headline misleads; with it, the honest comparisons are all within-coverage: zero-shot-unseen against LR-on-unseen (0.80 vs 0.0 — the gain), and size-matched seen against unseen (a ~2–5 point transfer penalty on the strong wordings, wider on weak ones).

Wording is a provider

The range across four wordings is 13.5 points on seen intents, 17.9 on unseen, 22.4 on safety test, 29.0 on safety shift — every one of them at or above the 5-point bar the preregistration set for “large”. The naive w0 wins or ties everywhere on intent, which should embarrass fancier phrasing: my hand-written descriptions lost to underscores-replaced-by-spaces. And wording moves ranking, not just accuracy — safety w2 has AUROC 0.39, below chance. GPT-3’s WiC result (arXiv:2005.14165, §3.7: 49.4% despite many phrasings) warned that wording effort does not always rescue a task; here wording effort sometimes breaks one. FLAN’s ablation (arXiv:2109.01652, §4.3) said the gains come from the instructions rather than the volume — our analogue in miniature: the descriptions dominate the embedding model, not the reverse. If you take one operational rule from this chapter, it is that rule: version your verbalizers like code, because they move numbers like code.

Safety and speed

Safety has no unseen split (binary), so it runs test plus shift only. The LR refit reproduces both frozen numbers exactly (0.8966 test, 0.5496 shift). Zero-shot test lands 0.47–0.70 across wordings, every paired interval excluding zero in LR’s favor; on shift the picture frays — w3 reaches 0.664 against LR’s 0.5496 with an interval entirely below zero, the chapter’s lone zero-shot win on trained territory, on 262 items, with one author’s wording. Report it, do not lean on it.

Speed, device-separated as required: per item, embed-sim takes 14.7 ms warm p50 on CPU (18.6 p95) against 6.1 ms (8.5 p95) on GPU; tuned LR answers in 0.35 ms (0.47 p95) on CPU — thirty to forty times faster on the same device, never compared across devices. Zero-shot’s tax is paid twice: in accuracy on seen labels, and in milliseconds everywhere.

Calibration is reported and disclaimed together: zero-shot ECE runs 0.37–0.59 at τ = 10.0 with AUROC 0.97–0.98 — it ranks excellently and calibrates terribly, exactly as declared. The temperature was arbitrary; Chapters 9–11 own what to do about it.

Against the prediction

Clause by clause: (a) seen loss supported, 17–30 points with intervals excluding zero on every wording; (b) unseen gain supported — 0.62–0.80 against 0.0 by construction, far above the 0.10 refutation floor; (c) wording supported — ranges of 13–29 points everywhere. The safety band prediction (0.60–0.75) half-held: w3 lands inside at 0.698, the rest below. The BTZSC-band expectation broke upward on unseen (0.80 vs 0.4–0.55) — reported, not hidden; crisp commercial intents plus aligned descriptions is the likely reason, and it is an interpretation, not a measurement.

Limitations

  • One embedding model. bge-small only; a bake-off is out of scope (BTZSC already ran it). A different encoder could shift every margin here.
  • Four wordings by one author. The measured spread is spread across my phrasings, with my vocabulary habits baked in — not across all possible wordings. The fixture records this; the numbers do not generalize past it.
  • τ = 10.0 is arbitrary. ECE reflects it; no calibration claim is made or implied.
  • The test inversion needs its control. Unseen-vs-seen test accuracy is confounded by set size (17 vs 60); the chapter’s comparisons are within-coverage, and the size control lives on threshold rows.
  • Safety is thin. n=116 test, n=262 shift; the w3 shift win is one wording on 262 items — an observation with an interval, not a finding to build on.
  • test opened once per task (this chapter’s own touch; touch-count rows in the file). All selection predates the run; nothing was chosen on test.
  • No live Jev behavior; no hosted calls. The Jev-vs-zero-shot comparison is future work.

What Chapter 5 leaves behind

  • src/arbiter/providers/embed_zero_shot.py — cosine-similarity zero-shot over bge-small-en-v1.5 (pinned snapshot), fixed τ, refusal on unknown candidate sets.
  • examples/ch05-zero-shot-decisions/ — frozen paraphrase fixtures (77×4 + 2×4, seeded split), runner, size-control analysis, selftest, demo, requirements, README.
  • results/ch05.jsonl — 1241 rows.
  • evidence/notes-ch05.md, research/ch05-additions.md, evidence/ledger.md (claims 5.1–5.19).

Is a runtime-defined label set a model property or an interface property? Both, split cleanly: the refusal mechanism is the interface property — it is what makes offering unseen labels safe, and any provider that guesses instead of refusing fails it. The ranking is the model property — embeddings supply it, with wording as its loudest knob. Chapter 6 asks whether entailment, the next family up in cost, changes either half.