← Jev From First Principles

where decide

Is semantic filtering a construct, or retrieval plus a decision?

where decide

This chapter reports a CPU experiment that ran here: a cross-encoder (ms-marco-MiniLM-L-6) and an encoder (bge-small) over SciFact. A cross-encoder is a relevance scorer, not a decision model in the Jev sense. Nothing here is evidence about Jev or about H1-H4.

The problem

Chapter 21 measured a pipeline where retrieval found the evidence and the decision layer threw a third of it away. The decision layer was a thresholded embedding cosine, and it was weak. So the question is not whether to filter, but which filter, and does the filter deserve syntax?

The construct under test is where:

docs.where(decide("contains evidence relevant to this claim"))

The chapter question is whether this is a construct or just retrieval plus a decision.

What we expect and why

Our starting hypothesis is that a cascade matches the cross-encoder’s quality at lower cost, and a cross-encoder matches both on simple relevance.

Three papers frame the comparison:

  1. Patel et al. (2024) — LOTUS’s sem_filter is a declarative operator: sem_filter(l: X -> Bool) returns the records that pass a natural-language predicate. Its gold algorithm is a model call per record; its optimization is a cheaper proxy (a small LLM, or embeddings) plus thresholds learned by sampling that meet recall/precision targets with probability 1 - delta. This is the closest published prior art, and it must be compared against honestly.

  2. Nogueira & Cho (2019) — a BERT cross-encoder re-ranks retrieved passages and is the top entry on MS MARCO at the time, beating the prior state of the art by 27% relative in MRR@10. The cross-encoder is the classical filter-after-retrieve baseline. (Read at abstract level.)

  3. Khattab & Zaharia (2020) — ColBERT’s late interaction encodes query and document separately and interacts cheaply, giving cross-encoder-like quality at far lower cost. The cheaper middle ground between bi-encoder and cross-encoder. (Read at abstract level.)

The build

src/arbiter/semantic_ops.py implements four filters over the same candidates.

from arbiter.semantic_ops import (
    apply_sigmoid_map,
    error_2x2,
    f1_threshold,
    fit_sigmoid_map,
    kept_metrics,
    precision_at_threshold,
)
    # 1. A threshold by pair-level F1, on hand labels.
    print("1. f1_threshold")
    scores = [0.1, 0.2, 0.8, 0.9]
    labels = [False, False, True, True]
    tau, f1, prec, rec = f1_threshold(scores, labels)
    print(f"   tau={tau} F1={f1:.2f} P={prec:.2f} R={rec:.2f}")
    # 2. A sigmoid calibration map fit on calibrate pairs.
    print("2. sigmoid calibration")
    cal_scores = [-1.0, 0.0, 1.0, 2.0, 3.0]
    cal_labels = [False, False, True, True, True]
    params = fit_sigmoid_map(cal_scores, cal_labels)
    probs = apply_sigmoid_map(cal_scores, params)
    print(f"   a={params[0]:.3f} b={params[1]:.3f}")
    print(f"   p(score=-1)={probs[0]:.3f}  p(score=3)={probs[-1]:.3f}")
    # 3. Kept-set precision/recall.
    print("3. kept_metrics")
    m = kept_metrics([True, True, False, False], [True, False, False, False])
    print(f"   precision={m['precision']:.2f} recall={m['recall']:.2f}")
    # 4. The query-level 2x2.
    print("4. error_2x2 (retrieved x kept)")
    x2 = error_2x2([True, True, False], [True, False, False])
    print(f"   retrieved_kept={x2['retrieved_kept']} retrieved_dropped={x2['retrieved_dropped']} "
          f"missed={x2['missed']}")
    print(f"   end_to_end={x2['end_to_end']:.3f}")
    # 5. Same candidates, different filters: which pairs each keeps.
    print("5. same candidates, four filters (hand-supplied)")
    emb = [0.80, 0.55, 0.40]
    ce = [4.0, -1.0, 2.5]
    labels = [True, False, True]
    tau_emb = 0.50
    tau_ce = 0.0
    emb_kept = [s >= tau_emb for s in emb]
    ce_kept = [s >= tau_ce for s in ce]
    print(f"   embedding keeps {sum(emb_kept)} of 3 (P/R: "
          f"{precision_at_threshold(emb, labels, tau_emb)[0]:.2f}/"
          f"{precision_at_threshold(emb, labels, tau_emb)[1]:.2f})")
    print(f"   cross-enc keeps {sum(ce_kept)} of 3 (P/R: "
          f"{precision_at_threshold(ce, labels, tau_ce)[0]:.2f}/"
          f"{precision_at_threshold(ce, labels, tau_ce)[1]:.2f})")

The walkthrough prints:

1. f1_threshold
   tau=0.8 F1=1.00 P=1.00 R=1.00
2. sigmoid calibration
   a=1.047 b=-0.440
   p(score=-1)=0.184  p(score=3)=0.937
3. kept_metrics
   precision=1.00 recall=0.50
4. error_2x2 (retrieved x kept)
   retrieved_kept=1 retrieved_dropped=1 missed=1
   end_to_end=0.333
5. same candidates, four filters (hand-supplied)
   embedding keeps 2 of 3 (P/R: 0.50/0.50)
   cross-enc keeps 2 of 3 (P/R: 1.00/1.00)

The experiment

python examples/ch22-where-decide/run_ch22.py (CPU, cached encoders, SciFact). All four filters see the same candidate pool: BM25 top-10 per claim, so 1,880 pairs on the test split. Thresholds come from calibrate (pair-level F1) except the decision filter, whose threshold comes from the Chapter 11 select_threshold on threshold (target selective risk 0.20, delta 0.10).

The result

Test split (results/ch22.jsonl, stage=filter), n=188 claims / 1,880 pairs:

Filter precision recall F1 kept end-to-end decision error CE passes
embedding 0.600 0.577 0.588 175 0.532 0.388 0
cross-encoder 0.553 0.654 0.599 215 0.606 0.314 1,880
decision (risk-controlled) 0.884 0.209 0.338 43 0.202 0.718 1,880
cascade (prefilter 5, then CE) 0.622 0.643 0.632 188 0.601 0.319 940

Cross-encoder latency: 24.70 ms per pair on this CPU (stage=latency).

What surprised us

  1. The cross-encoder helps, but only on recall. It lifts recall from 0.577 to 0.654 and end-to-end from 0.532 to 0.606, while its precision is lower (0.553 vs 0.600). A stronger relevance model did not cleanly dominate a cheap cosine. This is the Chapter 21 decision-error story continuing: no filter here is good.

  2. The cascade is the best operating point. It matches the cross-encoder’s end-to-end (0.601 vs 0.606) and its recall (0.643 vs 0.654) at half the forward passes (940 vs 1,880). That is the prediction confirmed, and it is what LOTUS’s proxy-plus-oracle design buys.

  3. The risk-controlled decision filter is too conservative to be useful. Configured with the Chapter 11 select_threshold at target risk 0.20 with delta 0.10, it accepted only 2.6% of the threshold pairs and kept just 43 of 1,880 test pairs. Its precision is the best (0.884) and its end-to-end is the worst (0.202). The distribution-free risk bound buys precision by spending almost all the coverage. P3 is refuted as stated: the calibrated, abstaining filter does not match the plain threshold; it is a precision-first point, not a dominant one.

  4. No filter keeps all retrieved evidence. Every filter’s decision error (0.314-0.718) exceeds retrieval error (0.080). P4 is supported, and it sharpens Chapter 21’s result: the decision stage, however implemented, is where the end-to-end loss lives on this task.

Prior art: where where decide stands against LOTUS

LOTUS’s sem_filter (arXiv:2407.11418) is the closest published operator. Stated exactly:

  • Same. Both filter records by a natural-language predicate. Both can use a cheap proxy first and a stronger model second. LOTUS’s sem_filter prototype is l_M(t_i) over each tuple; our where decide is that filter with a decision as the predicate.
  • Narrower. LOTUS’s sem_filter returns a Boolean per record. Our filter’s predicate is a typed decision: it can carry a calibrated probability and an abstention, and the abstention is exercised (the removed records are visible as abstained, not silently dropped). A Boolean cannot represent “I am not sure”.
  • Different. LOTUS specifies a gold algorithm per operator and provides statistical accuracy guarantees (recall and precision targets gamma_R, gamma_P met with probability 1 - delta) via sampling on the proxy, with a cost-based optimizer over an operator algebra. Our decision filter uses the Chapter 11 risk machinery instead of LOTUS’s sampling framework, and lives in a language, not a data-processing engine. LOTUS’s gold algorithm is an LLM call per tuple; ours is a CPU cross-encoder.

We claim no novelty. where decide is sem_filter with a typed, abstaining predicate and a different guarantee machinery. The chapter’s result is not that the construct is new; it is that, on this task, the filter choice matters more than the syntax, and that the risk-controlled predicate is the wrong operating point.

Wrong / Correct. Wrong: “where decide is a new semantic operator; the syntax is the contribution.” Correct: where decide is LOTUS’s sem_filter under another name. The contribution, if any, is the typed abstaining predicate and the calibration discipline — and on this task that discipline hurt.

The distinction this chapter keeps

A cross-encoder is a relevance scorer, not a decision model. It produces a ranking score. Calling its thresholded output a “decision” is a use of the contract, not evidence about a decision model. The typed decision (probability, abstention, refusal) is what the contract adds; the scorer is what the provider is.

Filter quality vs filter cost. The cascade is the Pareto point here: cross-encoder quality at half the passes. The risk-controlled filter is a different point (precision-first) and is not comparable on end-to-end accuracy alone.

What to carry forward

where decide is a method, and the library form is sem_filter plus a threshold. The construct earns syntax only if it enforces a discipline the library form cannot, and that discipline measurably prevents failure. The abstention discipline is the candidate; on this evidence it did not prevent failure, it caused it. That is a result for the language hypotheses (L1) and it is decided (if at all) in Chapter 35.

Chapter 23 asks the next question: if filtering is a method, what about iteration — is for decide a construct or sugar for filter-then-loop?

Close by

Does it deserve syntax, or is it a method? On this evidence, a method. The next chapter tests the same question for iteration.

Limitations

  • One dataset (SciFact), one encoder, one cross-encoder, one CPU machine. Not a claim about filters in general.
  • The decision filter’s conservatism is partly a configuration choice (target risk 0.20, delta 0.10). A higher risk target would trade precision for coverage; the chapter reports the configured point only.
  • The cross-encoder is used zero-shot, not fine-tuned on SciFact; a fine-tuned cross-encoder could change the ranking.
  • The cascade prefilter size m=5 is fixed, not tuned; the cascade’s cost saving is reported for that m.
  • LOTUS read in part (Abstract, §1, §2.1-§2.3, §3.1). Nogueira and ColBERT were read in full text after drafting (Session R); LOTUS §5 (evaluation) remains unread, and no claim here rests on it.
  • where decide is not implemented as syntax; the chapter tests the filter, and infers the construct question from it.