← Jev From First Principles

Abstain

A decision may refuse to decide. You can quote what refusing costs, and you must not trust the quote off-distribution.

Design draft: this chapter argues from the literature and from earlier measured results; its own experiment has not been run.

Your program branches on a provider’s score. Two chapters back, Confidence Is Not Probability showed you that the number is not a probability of being right, and Calibration gave you maps that fix that, inside the distribution they were fitted on. One option you have not yet given the program is the third answer: I will not act on this one. A decision that is never allowed to refuse must act on its worst guesses as surely as its best. Your queue of requests contains items your provider is confident and wrong about; calibration cannot remove them, it can only price them.

This chapter makes abstention first-class. Decision<T> = Decided<T> | Abstain, where Abstain carries a reason. Then it gives you the accounting: what coverage you keep when you buy a target error rate, and whether the quote survives when the input distribution changes. The short answer from the literature, and from the shift rows you already measured, is that the quote does not travel. The chapter’s job is to make both halves precise.

The origin still sets the frame

The reject option

DOCUMENTED: The reject option is old. C. K. Chow’s On optimum recognition error and reject tradeoff (IEEE Trans. Information Theory, 1970) is the standard origin citation for trading recognition error against rejection. We verified that the record exists via Crossref (DOI 10.1109/tit.1970.1054406) but could not open the full text behind the IEEE paywall, so we take the formal frame from the paper that made it operational for deep networks.

Selective classification: a pair, a curve, and a threshold

DOCUMENTED: Geifman and El-Yaniv define the selective classifier as a pair (f, g): a predictor plus a binary selection function. Coverage is φ = E[g(x)] and selective risk is R(f,g) = E[loss·g] / φ. The risk-coverage curve traces risk as a function of coverage. A risk target r* and confidence δ select a threshold over a validation sample so that R ≤ r* with probability at least 1−δ (their SGR procedure, Algorithm 1). [Geifman & El-Yaniv, arXiv:1705.08500, §§2, 3, 4]

Two details matter for us. First, their softmax-response (SR) score is max_j p_j, and they are explicit that it is a ranking: “the ideal confidence-rate function should only provide coherent ranking rather than absolute probability values.” A threshold on SR is a filter on an ordering, not a probability claim, which is exactly the distinction Chapter 9 taught you. Second, they add the caveat that if the score is severely skewed, “the bound of the resulting selective classifier can be far from the target risk.”

Under shift, the quote fails

DOCUMENTED, and it contradicts a comfortable reading of the above: Kamath, Jia and Liang test the same construction under domain shift and find the point threshold fails. Their base QA model is underconfident in-domain and overconfident out-of-domain: at a MaxProb of 0.6 the model is roughly 80% correct in-domain and 45% correct out-of-distribution. Mixing the two at test time, “MaxProb therefore does not abstain enough on the OOD examples.” [Kamath et al., arXiv:2006.09462, §§4.1, 5.2, 5.3]

Training a calibrator on a mixture of in-domain and known out-of-domain data raises coverage at 80% accuracy from 48.2% (MaxProb) to 56.1%, and their AUC numbers move from 20.5 and 19.3 (MaxProb variants) to 18.5 (calibrator). The direction that matters for us: a threshold picked in-domain cannot be assumed to hold elsewhere.

A 2026 paper that sharpens the caution

DOCUMENTED: Salem and colleagues measure that standard conformal risk control, the guaranteed-threshold method Chapter 10 built, “violates the per-group risk budget in up to 47% of evaluation trials” under mild shift in group composition, because the marginal guarantee is not a per-group one. Their hierarchical group-conditional fix (one threshold per node, Bonferroni-corrected, leaf-first selection) restores the guarantee at a cost of 22 to 37 percentage points of coverage on ARC-Challenge. [Salem et al., arXiv:2607.24562, §§1, abstract]

So the honest prediction is not simply “guarantees fail”. It is: a guarantee is a statement about the distribution you calibrated on; group-composition shift turns it into a statement about the average, and the average tolerates subgroups paying for the group.

What you have already measured

The book did not need an experiment to predict the failure; you have the rows.

OBSERVED: Chapter 10’s calibration run on safety. Temperature scaling reduces shift-minus-test median ECE by 0.3959 (tfidf-lr), 0.2107 (fasttext-style) and 0.1238 (embed-sim), and raw conformal median coverage on the shift split falls to 0.5687, 0.5611 and 0.8550, against 0.9088, 0.9107 and 0.9166 on intent test in-distribution. (Rows: results/ch10.jsonl, JEV-10-01, metric rows split: shift versus split: test, method temperature and conformal coverage; the P4 verdict was also REFUTED for safety embed at 0.8793, below its 0.88 floor, even with no shift.) The same distribution change that breaks calibration will break any threshold built on calibrated scores.

OBSERVED: The in-distribution floor from The Smallest Decision: tuned LR at 0.8779 on intent, fasttext-style at about 0.9052 on safety test. These become the no-abstention risk you are trying to buy down.

PROPOSED: The experiment below needs no new inference. Chapter 9 left the frozen per-item distributions (results/ch09.jsonl), Chapter 10 the fitted maps (evidence/ch10-fits/), and its handover records that the threshold split of each dataset was deliberately left untouched for the chapters after 10. That is the split this chapter’s experiment selects its threshold on. This is a property of the repository, not a measured result; it is stated so the experiment can be checked.

The distinction the chapter keeps

Three things look alike and are not:

  • Refusal: the provider cannot satisfy the request (answer set too large, no model loaded, wrong question type). arbiter.contract.Refusal, from Chapter 8. It happens before any decision; the caller cannot tune it away.
  • Abstain: the provider answered, and the caller’s policy declines to act on that answer. This chapter’s outcome. It happens after the distribution is on the table.
  • UNKNOWN / NONE_OF_THE_ABOVE: the question may be outside the options at all. That is Chapter 12’s territory, and the survey we read warns that answerability “is difficult to model in terms of model confidence”. [Wen et al., TACL 2025, §1]

Abstention is a caller-side action with a price, and the price is quoted in risk-coverage terms. A 2026 result adds the other half of the caution: the score you threshold matters as much as the threshold. DOCUMENTED: Phillips and colleagues find entropy-based uncertainty “model-dependent”, with a “confidently wrong” regime where low-entropy hallucinations break selective prediction at strict risk targets. Their acceptance rule is exactly ours (pick the largest threshold whose observed error among accepted answers is at most the target), and they warn that “strong AUROC does not imply reliable selective prediction at strict safety thresholds”. [Phillips et al., arXiv:2603.21172, §§1, 2.1]

Wrong: “A well-calibrated threshold is a guarantee; it holds wherever you deploy the provider.”

Correct: “A threshold selected on one split quotes coverage at a target risk for that distribution. Off it, the quote is what you measure next, and the measurements you already have say it will not hold.”

    flowchart TD
    A[request] --> B[provider returns distribution]
    B --> C{score >= theta?}
    C -- no --> D[Abstain reason]
    C -- yes --> E[Decided value]
    D --> F[escalate / log / human]
    E --> G[act on value]
  

The arithmetic, made executable

The arithmetic is model-free and unit-tested (tests/abstain/, 13 tests), in src/arbiter/abstain.py. Everything below is examples/ch11-abstain/walkthrough_ch11.py, which you can run as it stands. Its three ten-item splits are hand-written and small enough to check with a pencil, so the output shows the arithmetic and the caller’s obligation, and nothing about any provider.

The setup

from arbiter.abstain import (
    Abstain,
    AbstainReason,
    Resolved,
    apply_threshold,
    aurc,
    gate,
    risk_coverage,
    select_threshold,
)

# Calibration split: scores from a provider, and whether its answer was right.
CAL_SCORES = [0.95, 0.90, 0.85, 0.80, 0.70, 0.65, 0.60, 0.50, 0.40, 0.30]
CAL_CORRECT = [True, True, True, True, False, True, False, False, True, False]
# A fresh split from the same distribution, and one where errors sit higher up.
TEST_SCORES = [0.93, 0.88, 0.84, 0.78, 0.72, 0.66, 0.58, 0.47, 0.41, 0.28]
TEST_CORRECT = [True, True, True, False, True, True, False, False, True, False]
SHIFT_SCORES = [0.94, 0.91, 0.86, 0.82, 0.75, 0.69, 0.62, 0.55, 0.45, 0.35]
SHIFT_CORRECT = [True, False, True, False, False, True, False, True, False, False]


def main() -> None:

1. Risk and coverage

risk_coverage sorts items by score, highest first, and for every acceptance prefix measures the error rate among the accepted (selective risk, the Geifman definition). The table is the whole idea: accept the top four and you are never wrong, accept five and the first error appears. AURC is the trapezoidal area under this curve; lower is better.

    # 1. The risk-coverage table: accept the top k items, measure the error among them.
    print("1. risk and coverage, accepting the top k items on the calibration split")
    coverage, risk = risk_coverage(CAL_SCORES, CAL_CORRECT)
    print("    k   score  right   coverage   risk")
    for k, (score, ok, cov, r) in enumerate(zip(CAL_SCORES, CAL_CORRECT, coverage, risk), start=1):
        print(f"   {k:>2}   {score:.2f}  {'yes' if ok else 'NO ':<5}   {cov:>6.2f}   {r:.3f}")
    print(f"   AURC {aurc(coverage, risk):.3f}")

2. Buying a target error rate

select_threshold implements the “largest prefix whose observed risk is at most the target” rule. Passing delta swaps the point estimate for a Gascuel and Caraux binomial bound, which is the SGR variant.

    # 2. Buying a target error rate: the largest prefix whose risk is within the target.
    print("2. choosing a threshold for a target risk")
    for target in (0.10, 0.20, 0.01):
        threshold, accepted, cov, achieved = select_threshold(CAL_SCORES, CAL_CORRECT, target)
        print(f"   target {target:.2f} -> threshold {threshold:.2f}, accepts {accepted} of 10, coverage {cov:.2f}, risk {achieved:.3f}")
    print("   even a 0.01 target is 'met', because the top four items happen to be right: 0 errors in 4 is not a certificate")
    for target in (0.30, 0.45):
        threshold, accepted, cov, upper = select_threshold(CAL_SCORES, CAL_CORRECT, target, delta=0.1)
        if threshold is None:
            print(f"   target {target:.2f} with a 90% binomial bound -> unreachable from 10 items: refuse everything")
        else:
            print(f"   target {target:.2f} with a 90% binomial bound -> accepts {accepted}, upper bound {upper:.3f}")

Two things in the output are worth stopping on. A target of 0.01 is “met” only because the top four items happen to be right, and zero errors in four is not a certificate. The binomial bound sees that: at 90% confidence it cannot certify a 0.30 risk from ten items at all, so it refuses everything, and a 0.45 target is the smallest it will honour, accepting four items with an upper bound of 0.438. The point-estimate rule is what the literature calls the common practice; the bound is what a guarantee costs.

3. The threshold on a fresh split, and on a shifted one

apply_threshold is the only function the evaluation of test and shift runs, so nothing leaks between selection and evaluation.

    # 3. The same threshold, on a fresh split and on a shifted one.
    print("3. the threshold for target 0.20 travels badly")
    threshold, _, cov, achieved = select_threshold(CAL_SCORES, CAL_CORRECT, 0.20)
    print(f"   calibration: coverage {cov:.2f}, risk {achieved:.3f}")
    for name, scores, correct in (("fresh iid", TEST_SCORES, TEST_CORRECT), ("shifted", SHIFT_SCORES, SHIFT_CORRECT)):
        _, risk_here, cov_here = apply_threshold(scores, correct, threshold)
        print(f"   {name:<10}: coverage {cov_here:.2f}, risk {risk_here:.3f}")
    print("   the shifted split accepts the same share of items and is wrong three times as often among them")

The fresh split from the same distribution reproduces the calibration quote exactly (coverage 0.60, risk 0.167). The shifted split has the same coverage and three times the error among accepted items, because its errors sit above the threshold. That is the Kamath mechanism, shaped by hand-written data: a threshold tuned to one distribution cannot know which way a shift will push.

4. The caller’s obligation, and a trap in the first draft

Resolved carries exactly one of value or abstain, enforced when it is built, so a Resolved() with neither or both is a ValueError. There is no constructor path that silently drops the refusal, and the reason is an enum (LOW_CONFIDENCE, OUT_OF_SCOPE, CONFLICTING_EVIDENCE), so an Abstain cannot be a bare None.

The first draft of this chapter, and the library’s own docstring, showed the caller handling it like this:

case Resolved(value=v):     pay_out(v)
case Resolved(abstain=a):   escalate(a.reason)

That is wrong, and it is wrong in exactly the way the chapter exists to prevent. In Python’s pattern matching Resolved(value=v) matches every Resolved, because a pattern variable captures None as readily as a value. Written first, it sends every abstention down the pay-out arm with v = None. The safe order names the class and puts the refusal first. Both are in the output below, and the trap is pinned by a test so it cannot return unnoticed.

    # 4. The caller's obligation: the refusal arm cannot be forgotten.
    print("4. the caller must handle the refusal")
    below = gate(0.30, threshold, value="refund")
    above = gate(0.90, threshold, value="refund")

    def safe(outcome):
        match outcome:
            case Resolved(abstain=Abstain() as a):
                return f"escalate ({a.reason.value})"
            case Resolved(value=v):
                return f"pay out {v}"

    def unsafe(outcome):
        match outcome:
            case Resolved(value=v):
                return f"pay out {v}"
            case Resolved(abstain=a):
                return f"escalate ({a.reason.value})"

    print(f"   score 0.30 -> safe: {safe(below)}")
    print(f"   score 0.90 -> safe: {safe(above)}")
    print(f"   score 0.30 -> unsafe order: {unsafe(below)}   (the first arm captures None)")

PROPOSED: This runtime totality is the bridgehead of the language claim in The Decision Expression and if decide. The construct’s reason to exist is that a caller who may ignore the abstain outcome, will, and this very example is the evidence: a careful author wrote the trap into a chapter about not writing it. The enforcement has to become compile-time, not courtesy. If later evidence shows the library form is as safe, construct L1 is strengthened; this chapter only provides the shape.

5. A scoring bug the metric catches

    # 5. A scoring bug the metric catches.
    print("5. a sign error in the score")
    flipped = [1.0 - s for s in CAL_SCORES]
    c2, r2 = risk_coverage(flipped, CAL_CORRECT)
    print(f"   AURC correct score {aurc(coverage, risk):.3f}, sign-flipped score {aurc(c2, r2):.3f}  (lower is better)")

What it prints

1. risk and coverage, accepting the top k items on the calibration split
    k   score  right   coverage   risk
    1   0.95  yes       0.10   0.000
    2   0.90  yes       0.20   0.000
    3   0.85  yes       0.30   0.000
    4   0.80  yes       0.40   0.000
    5   0.70  NO        0.50   0.200
    6   0.65  yes       0.60   0.167
    7   0.60  NO        0.70   0.286
    8   0.50  NO        0.80   0.375
    9   0.40  yes       0.90   0.333
   10   0.30  NO        1.00   0.400
   AURC 0.156
2. choosing a threshold for a target risk
   target 0.10 -> threshold 0.80, accepts 4 of 10, coverage 0.40, risk 0.000
   target 0.20 -> threshold 0.65, accepts 6 of 10, coverage 0.60, risk 0.167
   target 0.01 -> threshold 0.80, accepts 4 of 10, coverage 0.40, risk 0.000
   even a 0.01 target is 'met', because the top four items happen to be right: 0 errors in 4 is not a certificate
   target 0.30 with a 90% binomial bound -> unreachable from 10 items: refuse everything
   target 0.45 with a 90% binomial bound -> accepts 4, upper bound 0.438
3. the threshold for target 0.20 travels badly
   calibration: coverage 0.60, risk 0.167
   fresh iid : coverage 0.60, risk 0.167
   shifted   : coverage 0.60, risk 0.500
   the shifted split accepts the same share of items and is wrong three times as often among them
4. the caller must handle the refusal
   score 0.30 -> safe: escalate (low_confidence)
   score 0.90 -> safe: pay out refund
   score 0.30 -> unsafe order: pay out None   (the first arm captures None)
5. a sign error in the score
   AURC correct score 0.156, sign-flipped score 0.640  (lower is better)

Read it in order.

  • Section 1: the first error arrives at the fifth item, so risk is 0 up to coverage 0.4, jumps to 0.200 at 0.5, and ends at 0.400. AURC is 0.156.
  • Section 2: targets 0.10 and 0.01 both stop at four items, and 0.20 stops at six. The bound turns “met” into “certified”, and a ten-item split cannot certify much.
  • Section 3: the same threshold at coverage 0.60 gives risk 0.167 on a fresh split and 0.500 on a shifted one.
  • Section 4: the safe order escalates a score of 0.30 and pays out at 0.90; the unsafe order pays out None at 0.30.
  • Section 5: flipping the score’s sign takes AURC from 0.156 to 0.640. A metric that cannot see a sign error cannot be trusted with a threshold.

The experiment, designed but not run

JEV-11-01 is preregistered with status: NOT_RUN; the full block is in metadata/11-chapter.yaml.

Question: at a target error rate, how much coverage does each calibrated provider keep, and does the threshold survive a shift?

  • Score functions: calibrated max-probability (apply the Chapter 10 fitted map to the Chapter 9 frozen distribution, then take p_argmax), raw max-probability (SR), and margin (p1 − p2).
  • Providers and splits: tfidf-lr, fasttext-style and embed-sim on intent and safety. Calibration maps are already fitted (Chapter 10, calibrate split); threshold θ is selected on the threshold split (untouched); evaluated on test and shift, one touch per task.
  • Targets and seeds: r* in {0.05, 0.10}, seeds {0, 1, 2} matching the Chapter 10 fits; realised risk reported with a paired bootstrap over items.
  • Baselines: no abstention (coverage 1.0, risk equal to the Chapter 4 provider error); raw max-probability threshold; margin threshold.

Predictions (HYPOTHESIS):

  • P1. In-distribution, the calibrated max-prob threshold realises selective risk within [r*, r* + 0.02] for at least 2 of 3 intent providers. (Reason: the fitted map makes the score a better ranking inside its own distribution; Geifman’s SR plus Chapter 10’s fitted maps.)
  • P2. On shift, the same threshold realises risk above r* for at least 2 of 3 intent providers. (Reason: the observed shift calibration collapse in Chapter 10’s rows above; Kamath §5.3.)
  • P3. AURC(calibrated) ≤ AURC(raw) for tfidf-lr on intent test.
  • P4. Margin matches raw max-prob in coverage at fixed risk within 2 points. (Reason: Geifman’s caution that SR has no privileged status among rankings.)
  • P5. Calibrated max-prob beats no-abstention on realised risk at equal coverage 0.9 for tfidf-lr intent test.

Refutation: P1 is refuted if no intent provider meets the in-distribution tolerance. P2 is refuted if every intent provider holds on shift, which is the surprising result and the one the design is built to be able to observe. Salem et al. give the follow-up if P2 holds only partially: group-conditional calibration (their HG-CRC) is the natural second experiment, and it is not preregistered here so it cannot be reported as predicted.

PENDING_RUN: result for JEV-11-01, P1–P5 Command: python examples/ch11-abstain/run_ch11.py --target 0.05 0.10 --seeds 0 1 2 Fills: results/ch11.jsonl, static/figures/ch11-risk-coverage-*.png Stage plan (each ≤15 min, CPU-only, reusing Ch 9 frozen distributions + Ch 10 maps): S1 assemble per-item scores per provider/seed; S2 risk-coverage + AURC on threshold/test/shift; S3 select θ on the threshold split; S4 realised-risk with paired bootstrap on test and shift; S5 write rows and figures.

What would change your mind

Two results decide this chapter’s thesis. If, against the rows you already have, the calibrated thresholds hold on shift for every provider, the “thresholds fail under shift” half is refuted and the chapter’s caution is reduced to a calibration story about your particular distributions. If, equally surprising, no calibrated provider meets the in-distribution tolerance, the “calibrated providers dominate” half is refuted, and the answer to the chapter’s question becomes that abstention machinery is worth nothing until the score is good. Neither is tuned away; both are written in the preregistration before any run.

No pivot proposal is made in this chapter: it designs and waits, and there is no new measured evidence on which to propose one. If the group-conditional follow-up is ever run, it must be preregistered before it becomes a result, not after.

Limitations

Everything that is not measured is said plainly. The experiment is NOT_RUN: no provider was re-evaluated, no results/ch11.jsonl exists, and every number in the predictions is HYPOTHESIS or OBSERVED-from-earlier-chapters, kept apart. The walkthrough’s splits are hand-written toys.

Chow (1970) is ABSTRACT_ONLY: its existence is verified, its text unopened, and it is used for origin framing only, with formal definitions taken from Geifman (FULL_TEXT). NLI and option-scored providers have no fitted maps yet, so they are absent from the design rather than silently excluded. The three score functions are a small family; margin thresholds have known ties, and conflation of tied scores is not modelled here.

The typed enforcement is runtime-total, not compile-time; the language chapters own that claim. Legacy providers return Decision, not DecisionV3, and gating is defined on scores, so the calibration-state binding from Chapter 10 is respected only if the caller uses the v3 view. Safety n is small (test 116, shift 262), so the paired intervals will be wide, and the chapter reports them wide.

What Chapter 12 inherits

Abstain exists in the type system of this book, with a reason, a price in coverage, and a warning that the price does not travel. The open question is whether you can tell I cannot tell from there is nothing in the options to tell, with the survey’s warning in hand: answerability “is difficult to model in terms of model confidence”.

Chapter 12 (Unknown) takes the three outcomes (ABSTAIN, NONE_OF_THE_ABOVE, UNKNOWN) and asks whether any provider can be made to tell them apart. It inherits from here the caller-side gate it can repoint at a third reason, and the reading list that says the answer is not going to fall out of a confidence score.