Unknown
Separate 'I cannot tell' from 'none of these answers' and 'this question has no answer here' before asking whether providers can separate them.
Design draft: this chapter argues from the literature and from earlier measured results; its own experiment has not been run.
You have taught a program to refuse. Abstain gave the caller a typed way to decline a provider’s answer and a way to price that refusal. But one refusal shape is not enough. Consider three failures in a banking assistant. It sees a vague request to change a card limit and cannot choose between two nearly identical card intents. It sees a request to order paper checks when its menu contains only electronic banking actions. And it sees a refund instruction whose order ID is blank. The first is uncertainty among answers. The second is certainty that the answers are wrong. The third is not a decision at all: the mandatory state needed to decide is absent.
This chapter gives those three cases different names and different runtime duties:
- ABSTAIN: “I cannot tell which listed option is right.”
- NONE_OF_THE_ABOVE: “I can tell that none of the listed options is right.”
- UNKNOWN: “This question cannot be answered from this state.”
PROPOSED: Those are not three intensities of doubt. They are three claims about different objects: the evidence for options, the option set itself, and the adequacy of the state. A single confidence cutoff cannot establish all three, because the cutoff only measures the first.
What the older literature already separates
The unknown-label problem predates decision models, and the useful papers separate scoring from naming.
Maximum softmax probability: a baseline, with a warning attached
DOCUMENTED: Hendrycks and Gimpel define both error detection and in/out-of-distribution detection, then use AUROC because accuracy can be misleading when base rates differ. Their baseline retrieves the maximum softmax probability: correctly classified examples tend to have larger maxima than incorrect and out-of-distribution examples. [Hendrycks & Gimpel, arXiv:1610.02136, Abstract, §§1–2]
They immediately add the warning that matters here: softmax probability has “a poor direct correspondence to confidence,” including a case where random Gaussian noise receives 91% MNIST “prediction confidence.” The baseline is useful for ranking and not trustworthy as a confession.
Their text-categorisation experiment is the closest analogue to BANKING77. They train on a subset of subjects (15 of 20 Newsgroups, 6 of Reuters 8, 40 of Reuters 52) and treat the held-out subjects as out-of-distribution. Maximum softmax probability distinguishes in- from out-of-distribution subjects with AUROC 0.75, 0.92 and 0.95 respectively. That is genuine support for the score-based half of our hypothesis, but it is not a transferable constant: those are different models, different held-out subjects, and in one case only two held-out topics. [Hendrycks & Gimpel, §3.2.2, Table 6]
Open-set recognition: naming the unknown is a different problem
DOCUMENTED: Bendale and Boult make the conceptual move this chapter needs. A closed-set network “forces” every input into a known class. One may train an “other” class for known unknowns, but it is impossible to train on every unknown unknown. Thresholding softmax probabilities helps but does not satisfy their definition of open-set recognition. [Bendale & Boult, arXiv:1511.06233, Abstract, §§1–3]
Their OpenMax alternative estimates the probability of an input being unknown from penultimate-layer activation vectors, using extreme-value-theory Weibull models and a compact abating probability construction. On their 80,000-image test, OpenMax improves F-measure by nearly 4.3 percentage points over optimally thresholded softmax and 12.3 points over the base network: 3,450 and 9,847 more correct outcomes respectively.
The same paper limits its own victory in ways our design copies. Rejected inputs still need an operational owner, which the authors leave to the system designer. Their reported comparison varies the shared uncertainty threshold, while wider OpenMax parameter sensitivity lives in the supplement. Larger tail sizes reject more unknowns but also more true images, so tail size 20 is an optimum to preserve, not a universal law. [Bendale & Boult, §§4, 6.1]
Energy: a score that keeps what softmax throws away
DOCUMENTED: Liu and colleagues give energy its exact meaning. For logits f(x) and temperature T, free energy is E(x; f) = −T · log Σᵢ exp(fᵢ / T). Lower energy means in-distribution and higher means out-of-distribution. At inference time they use negative energy as the OOD score and choose its threshold from in-distribution data. [Liu et al., arXiv:2010.03759, Abstract, §§2–3.1, 4.2]
On CIFAR-10 with WideResNet, energy reduces average FPR at 95% true-positive rate by 18.03% relative to softmax confidence. On CIFAR-100 the corresponding reduction is smaller, while AUROC rises from 0.9090 to 0.9188 on CIFAR-10 and from 0.7553 to 0.7956 on CIFAR-100. Energy is parameter-free at inference, but fine-tuning with auxiliary outliers can widen the gap further.
Two newer intent studies, and one that contradicts the universal form
Those three papers fix the methods. Three newer studies decide how narrowly we may state the “explicit other” prediction.
DOCUMENTED: Wang and colleagues prompt ChatGPT for out-of-domain intent detection with an explicit unknown choice. On Banking-50%, their zero-shot OOD recall is 56 points and OOD F1 47.41 points below UniNL. On full CLINC, with 150 intents in context, ChatGPT returns neither an in-domain intent nor “unknown” on approximately 8.49% of test samples. Longer intent lists make the model more likely to misclassify OOD items as in-domain. [Wang et al., arXiv:2402.17256, §§4.2, 5.2]
This supports brittleness, but only for the prompted-LLM mechanism: an “unknown” string inside a long instruction is not the same object as a scored NOTA label in a probability provider.
DOCUMENTED: Liu and colleagues’ LLM study qualifies the score half. A plain cosine-distance detector performs exceptionally well because LLM embedding spaces are comparatively isotropic, unlike the narrow-cone BERT-family spaces. LLMs are natural far-OOD detectors, while near-OOD (the same-domain, different-intent case closest to BANKING77) needs either a 65-billion-parameter model or fine-tuning. Score choice is therefore an empirical question, not a corollary of calibration. [Liu et al., arXiv:2308.10261, §1]
DOCUMENTED, and it contradicts the universal form of our hypothesis: Sali and Toraman implement OOD as an explicit tool named “fallback,” enforce tool calling, and ask the model to select among intent tools plus fallback. Their hybrid method then uses a small classifier for in-domain prediction and asks GPT-4o to verify or reject that prediction. Across six near-OOD datasets, the highest OOD F1 in every dataset belongs to a hybrid method. [Sali & Toraman, Findings of EMNLP 2025, §§3.3.2, 3.3.4, 4.1]
Forcing the output shape and narrowing the LLM’s task changes the result: explicit “other” is brittle as an appended probability label, but workable as an enforced tool outcome plus verification.
Wrong: “Add NONE_OF_THE_ABOVE to every option list; unknown handling is then solved.”
Correct: “An explicit other label is one mechanism. A threshold is another. Missing state is neither. The experiment must compare mechanisms, not names.”
flowchart TD
A[state + question] --> B{mandatory state present?}
B -- no --> C[UNKNOWN]
B -- yes --> D[provider returns distribution]
D --> E{explicit NOTA selected?}
E -- yes --> F[NONE_OF_THE_ABOVE]
E -- no --> G{score below cutoff?}
G -- yes --> H[ABSTAIN]
G -- no --> I[decide listed label]
What earlier chapters already measured
The design reuses a frozen unknown pool; it does not invent one.
OBSERVED: Zero-Shot Decisions froze BANKING77 into 60 seen labels with 2,400 test items and 17 unseen labels with 680 test items. With the naive wording, embedding similarity reaches 0.7312 on seen and 0.7882 on unseen. The headline inversion is a candidate-set-size effect: at matched 17-way size, the seen subset beats unseen on every wording. The tuned LR refuses all 680 unseen items when they are requested as unseen labels, 0.0 by construction, counted rather than smoothed. (Rows: results/ch05.jsonl, provider=embed-sim, dataset=banking77-seen/banking77-unseen, split=test, metric=accuracy; LR refusal in the same file and Chapter 5 §“The gain on unseen labels.”)
That refusal is important and easy to misread. It is provider-level refusal: the requested option labels were outside the fitted label set. It does not show that LR recognised an unknown banking request. PROPOSED: JEV-12-01 therefore asks unknown texts about the 60 seen options. That is the open-set question; Chapter 5’s refusal row is the interface property that made the old question safe, not the semantic answer to the new one.
OBSERVED: Decisions as Entailment warns against treating every low score as unknown. On safety shift, both NLI models rank items near or below chance while still answering: DeBERTa T3 AUROC 0.3775 and BART T3 0.3044 on 262 shift items. (Rows: results/ch06.jsonl, split=shift, metric=auroc.) A detector can fail while a classifier keeps answering. The same failure shape is why this chapter needs both ranking metrics and outcome confusion, not ranking alone.
The First Token adds a presentation warning without needing a new number here: changing the option list can change the readout problem. Appending NOTA is therefore an intervention on the measurement instrument, not a neutral extra row. The contract records that intervention through presentation evidence; the experiment reports NOTA selection on known items alongside unknown detection.
The outcome algebra, made executable
The model-free core is src/arbiter/unknown.py, exercised by 12 tests in tests/unknown/. It does not contain an OOD detector. It contains the outcome algebra an OOD experiment needs: the three distinct outcomes, the order in which they are decided, and the helpers that score and tabulate them. Everything below is examples/ch12-unknown/walkthrough_ch12.py, which you can run as it stands. Every score and logit in it is hand-supplied and small enough to check by hand, so the output shows how the helpers behave and nothing about any provider.
The setup
from math import exp
from pathlib import Path
import sys
ROOT = Path(__file__).resolve().parents[2]
sys.path[:0] = [str(ROOT), str(ROOT / "src")]
from arbiter.contract import ChoiceAnswer
from arbiter.unknown import (
NONE_OF_THE_ABOVE_LABEL,
UNKNOWN_LABEL,
confusion_counts,
free_energy,
known_unknown_auroc,
outcome_from_choice,
resolve_open_set,
)
def softmax(logits):
peak = max(logits)
exps = [exp(x - peak) for x in logits]
return [e / sum(exps) for e in exps]
def main() -> None:
1. The three situations, as three calls
A vague request becomes an Abstain through the score cutoff. A NONE_OF_THE_ABOVE winner becomes a NoneOfTheAbove that records exactly which listed options it rejected, and it refuses to exist when there were no listed options to reject. Unknown cannot be built from an answer alone: the caller must say which mandatory fields are absent.
# 1. The three situations from the chapter's opening, as three calls.
print("1. three ways of not deciding")
vague = ChoiceAnswer("limit_up", {"limit_up": 0.40, "limit_down": 0.38, "other": 0.22}, 0.02)
out = resolve_open_set(vague, abstain_below=0.60)
print(f" vague card request -> {type(out).__name__}: reason={out.reason.value!r}")
nota = ChoiceAnswer(
NONE_OF_THE_ABOVE_LABEL,
{"transfer": 0.15, "card_block": 0.10, NONE_OF_THE_ABOVE_LABEL: 0.75},
0.50,
)
out = outcome_from_choice(nota)
print(f" order paper checks -> {type(out).__name__}: rejected={list(out.rejected_options)}")
blank = ChoiceAnswer("refund", {"refund": 0.95, "exchange": 0.05}, 0.90)
out = resolve_open_set(blank, abstain_below=0.60, missing=["order_id"])
print(f" blank order id -> {type(out).__name__}: missing={list(out.missing)}")
2. Precedence and the cutoff
resolve_open_set gives missing state precedence over every score, so a confident 0.95 never gets to speak when the order ID is blank. It preserves explicit NOTA and ABSTAIN winners, and otherwise applies one inclusive cutoff. abstain_below=None disables score rejection explicitly, so “no threshold” is a declared policy and not an omitted argument.
# 2. Precedence and the cutoff.
print("2. precedence and the cutoff")
precedence = resolve_open_set(blank, abstain_below=0.60, missing=["order_id"])
print(f" confident 0.95 refund, order id missing -> {type(precedence).__name__} (the score is never consulted)")
edge = ChoiceAnswer("refund", {"refund": 0.60, "exchange": 0.40}, 0.20)
print(f" score 0.60, cutoff 0.60 -> {resolve_open_set(edge, abstain_below=0.60)!r} (the cutoff is inclusive)")
print(f" score 0.60, cutoff 0.61 -> {type(resolve_open_set(edge, abstain_below=0.61)).__name__}")
print(f" score 0.60, cutoff None -> {resolve_open_set(edge, abstain_below=None)!r} (no threshold is a declared policy)")
3. Why energy needs logits and not probabilities
free_energy implements Liu’s definition from supplied logits. It does not accept a probability distribution as a substitute, because softmax forgets the additive logit offset and a Chapter 10 calibration map need not be invertible. Two logit vectors that differ by a constant have identical softmax and different energy. The future runner must therefore record raw logits (LR decision-function values, fastText final-layer values, temperature-scaled cosine similarities or NLI entailment logits) and not reconstruct them afterward.
# 3. Why energy needs logits and not probabilities.
print("3. softmax forgets the logit offset")
near = [3.0, 0.0, 0.0]
far = [13.0, 10.0, 10.0] # the same vector plus 10
print(f" max probability: {max(softmax(near)):.4f} vs {max(softmax(far)):.4f} (identical)")
print(f" free energy: {free_energy(near):.4f} vs {free_energy(far):.4f} (differ by exactly 10)")
4. The ranking wrapper
known_unknown_auroc reuses the harness’s AUROC with “known” as the positive class; it implements no new ranking statistic. With four known and four unknown scores you can count the pairs yourself: fourteen of sixteen known-unknown pairs rank correctly, which is 0.875. Flip the arguments and the same code gives 0.125, so a sign error cannot hide.
# 4. The ranking wrapper, small enough to check by hand.
print("4. known-versus-unknown AUROC")
known = [0.90, 0.80, 0.70, 0.50]
unknown = [0.60, 0.55, 0.40, 0.30]
print(f" known {known}, unknown {unknown}")
print(f" AUROC {known_unknown_auroc(known, unknown):.3f} (14 of 16 known-unknown pairs rank correctly: 7/8)")
print(f" scores flipped: AUROC {known_unknown_auroc(unknown, known):.3f} (a sign error reverses it)")
5. What the score-only picture hides
Open-set evaluation is rectangular. Truth distinguishes known, unknown-label and unanswerable-state items, while predictions distinguish labels and the three non-answer outcomes. In this toy table a score cutoff sends three of the four unknown items to abstain; only the NOTA label says “none of these”.
# 5. Outcomes are not interchangeable, so the matrix is rectangular.
print("5. what the score-only picture hides")
actual = ["known"] * 4 + ["unknown"] * 4 + ["unanswerable"] * 2
predicted = (["label", "label", "label", "abstain"]
+ ["abstain", "abstain", "abstain", "none_of_the_above"]
+ ["unknown", "abstain"])
counts = confusion_counts(
actual, predicted,
["known", "unknown", "unanswerable"],
["label", "abstain", "none_of_the_above", "unknown"],
)
header = ["label", "abstain", "none_of_the_above", "unknown"]
print(f" {'actual':<13}" + "".join(f"{h:>19}" for h in header))
for a, row in counts.items():
print(f" {a:<13}" + "".join(f"{row[h]:>19}" for h in header))
print(" a score cutoff sends most unknown items to abstain; only the NOTA label names 'none of these'")
print(f" (reserved labels: {NONE_OF_THE_ABOVE_LABEL}, {UNKNOWN_LABEL})")
What it prints
1. three ways of not deciding
vague card request -> Abstain: reason='low_confidence'
order paper checks -> NoneOfTheAbove: rejected=['transfer', 'card_block']
blank order id -> Unknown: missing=['order_id']
2. precedence and the cutoff
confident 0.95 refund, order id missing -> Unknown (the score is never consulted)
score 0.60, cutoff 0.60 -> 'refund' (the cutoff is inclusive)
score 0.60, cutoff 0.61 -> Abstain
score 0.60, cutoff None -> 'refund' (no threshold is a declared policy)
3. softmax forgets the logit offset
max probability: 0.9094 vs 0.9094 (identical)
free energy: -3.0949 vs -13.0949 (differ by exactly 10)
4. known-versus-unknown AUROC
known [0.9, 0.8, 0.7, 0.5], unknown [0.6, 0.55, 0.4, 0.3]
AUROC 0.875 (14 of 16 known-unknown pairs rank correctly: 7/8)
scores flipped: AUROC 0.125 (a sign error reverses it)
5. what the score-only picture hides
actual label abstain none_of_the_above unknown
known 3 1 0 0
unknown 0 3 1 0
unanswerable 0 1 0 1
a score cutoff sends most unknown items to abstain; only the NOTA label names 'none of these'
(reserved labels: NONE_OF_THE_ABOVE, UNKNOWN)
Those numbers are properties of hand-built inputs. The repository’s earlier demo used hand-separated synthetic pools whose AUROCs came out near-perfect, which taught nothing; the inputs here are small and imperfect on purpose so every figure can be checked by hand.
The experiment, designed but not run
JEV-12-01 is preregistered with status: NOT_RUN; the full block is in metadata/12-chapter.yaml.
Question: can any provider reliably separate “outside the options” from “unsure inside them”?
- Pools: frozen Chapter 5 seen-60 test items as known (n=2,400) and unseen-17 test items as unknown (n=680). Unknown texts are scored against the 60 seen options. Missing-state UNKNOWN cases use separate hand fixtures because BANKING77 does not encode required-slot preconditions.
- Literal NOTA: one frozen option,
NONE_OF_THE_ABOVE(“none of the listed intents”; NLI hypothesis: “This text is about none of the listed intents”). It applies to embedding and NLI providers. The trained closed classifiers cannot fit an unseen extra class, so that arm isNOT_APPLICABLEto TF-IDF+LR and fastText-style by construction. - Scores: raw max-probability and T=1.0 free energy from recorded non-probability logits. No calibrated map is inverted into logits.
- Operating points: score thresholds are matched to the literal-NOTA known false-non-answer rate on the known threshold split. Test pools are evaluated once.
- Providers: TF-IDF+LR (C=1000), fastText-style (Chapter 4 rev1), embedding similarity (pinned bge-small snapshot), DeBERTa-v3-large NLI and BART-large-mnli (Chapter 6 revisions and T1 to T3 templates, primary rule).
- Metrics: known-versus-unknown AUROC; open-set accuracy; the full rectangular outcome-confusion matrix. Bootstrap B=1000, seed 42.
Predictions (HYPOTHESIS):
- P1. Raw max-probability reaches known-versus-unknown AUROC of at least 0.70 for every evaluated provider. The reason is Hendrycks §3.2.2, tempered by Liu’s warning that near-OOD is the hard case.
- P2. Literal NOTA is unstable: for each applicable provider, either unknown NOTA recall trails the matched score-only unknown recall by at least 10 points, or known false-NOTA selection reaches at least 5%. The reason is Wang §4.2, scoped to direct label-append scoring; Sali §3.3 to §4.1 is the named exception mechanism.
- P3. LR energy is non-inferior to LR max-probability: the AUROC difference has a positive point estimate and a 95% lower bound above −0.02. The reason is Liu §4.2.
- P4. Score-only evaluation conflates its non-answers: at the matched operating point, unknown items assigned ABSTAIN outnumber unknown NOTA wins by at least two to one. The reason is that a low-score cutoff observes weak support, not the reason for weakness. Section 5 above is the toy version of this prediction.
Refutation: P1 fails on any provider whose interval lies entirely below 0.70. P2 fails if literal NOTA both detects unknown better and stays below the known-error bound. P3 fails if LR energy is worse beyond tolerance. P4 fails if NOTA wins at least as often as ABSTAIN on unknown items. The prompt permits the collapsed recommendation, keeping fewer than three outcomes, but only as a verdict after this matrix and not as a design preference before it.
PENDING_RUN: result for JEV-12-01, P1–P4 Command:
python examples/ch12-unknown/run_ch12.py --providers tfidf-lr,fasttext-style,embed-sim,nli-deberta,nli-bart --templates T1,T2,T3 --threshold-source threshold --bootstrap-seed 42Fills:results/ch12.jsonl,static/figures/ch12-open-set-*.pngStage plan (each ≤15 min; no downloads, hosted calls, retraining, or pre-run rows): S1 build frozen known/unknown/NOTA requests; S2 run providers once per pool and record raw logits; S3 match operating points and compute AUROC/intervals; S4 report open-set accuracy and outcome confusion; S5 write rows and figures.
What would change your mind
Three measured results would overturn the chapter’s working thesis.
- If literal NOTA wins on unknown items without taxing known accuracy, P2 is refuted and the “brittle” adjective is withdrawn for direct scoring, not merely qualified by Sali’s tool mechanism.
- If raw max-probability cannot separate held-out banking intents from seen ones, P1 is refuted and the Hendrycks text result does not carry over to this near-OOD task.
- If LR energy loses to max-probability beyond tolerance, P3 is refuted and Liu’s image-classifier ordering does not carry over to this linear text classifier.
No pivot proposal is made: a design contains no new measured evidence on which to propose one.
Limitations
The experiment is NOT_RUN: no provider was queried, results/ch12.jsonl does not exist, and every trial number above is HYPOTHESIS. Required papers are PARTIAL by section, not full-text audits; supplements, proofs and repository code were not reproduced. Sali and Toraman was read from the ACL PDF, not merely the abstract. Wang’s numbers concern prompted ChatGPT and GPT-4, not our providers.
Liu 2024’s cosine-distance finding is not a preregistered arm because the design compares the prompt’s three baselines; adding cosine distance later would be a dated amendment, not a silent extension. The NLI and embedding energy arms depend on recording raw logits at inference time; that instrumentation is designed, not implemented. LR and fastText have no literal-NOTA arm because a closed fitted label set cannot honestly contain an unfitted class.
Missing-state UNKNOWN is tested on hand fixtures, so it validates the representation and resolver, not semantic detection of unanswerability. Safety has no unseen-label split, so the unknown evaluation is intent-only. Test pools are evaluated once; threshold matching uses only known threshold data.
What the next chapters inherit
Chapter 13 inherits the outcome vocabulary its probes must eventually emit: a hidden-state head that only returns labels has not solved unknown handling.
Chapters 16 to 19 inherit the enforcement question: three distinct runtime outcomes are only as strong as the caller discipline that handles them. If JEV-12-01 finds the outcomes inseparable in practice, the language chapters inherit a collapse recommendation with the confusion matrix as its evidence.
The prompt’s closing question therefore stays open by design: should the language keep three outcomes the models cannot tell apart? This chapter answers that no model has yet been asked properly, and specifies the asking.