where decide
Is semantic filtering a construct, or retrieval plus a decision?
where decide
This chapter reports a CPU experiment that ran here: a cross-encoder (
ms-marco-MiniLM-L-6) and an encoder (bge-small) over SciFact. A cross-encoder is a relevance scorer, not a decision model in the Jev sense. Nothing here is evidence about Jev or about H1-H4.
The problem
Chapter 21 measured a pipeline where retrieval found the evidence and the decision layer threw a third of it away. The decision layer was a thresholded embedding cosine, and it was weak. So the question is not whether to filter, but which filter, and does the filter deserve syntax?
The construct under test is where:
docs.where(decide("contains evidence relevant to this claim"))
The chapter question is whether this is a construct or just retrieval plus a decision.
What we expect and why
Our starting hypothesis is that a cascade matches the cross-encoder’s quality at lower cost, and a cross-encoder matches both on simple relevance.
Three papers frame the comparison:
-
Patel et al. (2024) — LOTUS’s
sem_filteris a declarative operator:sem_filter(l: X -> Bool)returns the records that pass a natural-language predicate. Its gold algorithm is a model call per record; its optimization is a cheaper proxy (a small LLM, or embeddings) plus thresholds learned by sampling that meet recall/precision targets with probability1 - delta. This is the closest published prior art, and it must be compared against honestly. -
Nogueira & Cho (2019) — a BERT cross-encoder re-ranks retrieved passages and is the top entry on MS MARCO at the time, beating the prior state of the art by 27% relative in MRR@10. The cross-encoder is the classical filter-after-retrieve baseline. (Read at abstract level.)
-
Khattab & Zaharia (2020) — ColBERT’s late interaction encodes query and document separately and interacts cheaply, giving cross-encoder-like quality at far lower cost. The cheaper middle ground between bi-encoder and cross-encoder. (Read at abstract level.)
The build
src/arbiter/semantic_ops.py implements four filters over the same candidates.
from arbiter.semantic_ops import (
apply_sigmoid_map,
error_2x2,
f1_threshold,
fit_sigmoid_map,
kept_metrics,
precision_at_threshold,
)
# 1. A threshold by pair-level F1, on hand labels.
print("1. f1_threshold")
scores = [0.1, 0.2, 0.8, 0.9]
labels = [False, False, True, True]
tau, f1, prec, rec = f1_threshold(scores, labels)
print(f" tau={tau} F1={f1:.2f} P={prec:.2f} R={rec:.2f}")
# 2. A sigmoid calibration map fit on calibrate pairs.
print("2. sigmoid calibration")
cal_scores = [-1.0, 0.0, 1.0, 2.0, 3.0]
cal_labels = [False, False, True, True, True]
params = fit_sigmoid_map(cal_scores, cal_labels)
probs = apply_sigmoid_map(cal_scores, params)
print(f" a={params[0]:.3f} b={params[1]:.3f}")
print(f" p(score=-1)={probs[0]:.3f} p(score=3)={probs[-1]:.3f}")
# 3. Kept-set precision/recall.
print("3. kept_metrics")
m = kept_metrics([True, True, False, False], [True, False, False, False])
print(f" precision={m['precision']:.2f} recall={m['recall']:.2f}")
# 4. The query-level 2x2.
print("4. error_2x2 (retrieved x kept)")
x2 = error_2x2([True, True, False], [True, False, False])
print(f" retrieved_kept={x2['retrieved_kept']} retrieved_dropped={x2['retrieved_dropped']} "
f"missed={x2['missed']}")
print(f" end_to_end={x2['end_to_end']:.3f}")
# 5. Same candidates, different filters: which pairs each keeps.
print("5. same candidates, four filters (hand-supplied)")
emb = [0.80, 0.55, 0.40]
ce = [4.0, -1.0, 2.5]
labels = [True, False, True]
tau_emb = 0.50
tau_ce = 0.0
emb_kept = [s >= tau_emb for s in emb]
ce_kept = [s >= tau_ce for s in ce]
print(f" embedding keeps {sum(emb_kept)} of 3 (P/R: "
f"{precision_at_threshold(emb, labels, tau_emb)[0]:.2f}/"
f"{precision_at_threshold(emb, labels, tau_emb)[1]:.2f})")
print(f" cross-enc keeps {sum(ce_kept)} of 3 (P/R: "
f"{precision_at_threshold(ce, labels, tau_ce)[0]:.2f}/"
f"{precision_at_threshold(ce, labels, tau_ce)[1]:.2f})")
The walkthrough prints:
1. f1_threshold
tau=0.8 F1=1.00 P=1.00 R=1.00
2. sigmoid calibration
a=1.047 b=-0.440
p(score=-1)=0.184 p(score=3)=0.937
3. kept_metrics
precision=1.00 recall=0.50
4. error_2x2 (retrieved x kept)
retrieved_kept=1 retrieved_dropped=1 missed=1
end_to_end=0.333
5. same candidates, four filters (hand-supplied)
embedding keeps 2 of 3 (P/R: 0.50/0.50)
cross-enc keeps 2 of 3 (P/R: 1.00/1.00)
The experiment
python examples/ch22-where-decide/run_ch22.py (CPU, cached encoders, SciFact). All four filters see the same candidate pool: BM25 top-10 per claim, so 1,880 pairs on the test split. Thresholds come from calibrate (pair-level F1) except the decision filter, whose threshold comes from the Chapter 11 select_threshold on threshold (target selective risk 0.20, delta 0.10).
The result
Test split (results/ch22.jsonl, stage=filter), n=188 claims / 1,880 pairs:
| Filter | precision | recall | F1 | kept | end-to-end | decision error | CE passes |
|---|---|---|---|---|---|---|---|
| embedding | 0.600 | 0.577 | 0.588 | 175 | 0.532 | 0.388 | 0 |
| cross-encoder | 0.553 | 0.654 | 0.599 | 215 | 0.606 | 0.314 | 1,880 |
| decision (risk-controlled) | 0.884 | 0.209 | 0.338 | 43 | 0.202 | 0.718 | 1,880 |
| cascade (prefilter 5, then CE) | 0.622 | 0.643 | 0.632 | 188 | 0.601 | 0.319 | 940 |
Cross-encoder latency: 24.70 ms per pair on this CPU (stage=latency).
What surprised us
-
The cross-encoder helps, but only on recall. It lifts recall from 0.577 to 0.654 and end-to-end from 0.532 to 0.606, while its precision is lower (0.553 vs 0.600). A stronger relevance model did not cleanly dominate a cheap cosine. This is the Chapter 21 decision-error story continuing: no filter here is good.
-
The cascade is the best operating point. It matches the cross-encoder’s end-to-end (0.601 vs 0.606) and its recall (0.643 vs 0.654) at half the forward passes (940 vs 1,880). That is the prediction confirmed, and it is what LOTUS’s proxy-plus-oracle design buys.
-
The risk-controlled decision filter is too conservative to be useful. Configured with the Chapter 11
select_thresholdat target risk 0.20 with delta 0.10, it accepted only 2.6% of thethresholdpairs and kept just 43 of 1,880 test pairs. Its precision is the best (0.884) and its end-to-end is the worst (0.202). The distribution-free risk bound buys precision by spending almost all the coverage. P3 is refuted as stated: the calibrated, abstaining filter does not match the plain threshold; it is a precision-first point, not a dominant one. -
No filter keeps all retrieved evidence. Every filter’s decision error (0.314-0.718) exceeds retrieval error (0.080). P4 is supported, and it sharpens Chapter 21’s result: the decision stage, however implemented, is where the end-to-end loss lives on this task.
Prior art: where where decide stands against LOTUS
LOTUS’s sem_filter (arXiv:2407.11418) is the closest published operator. Stated exactly:
- Same. Both filter records by a natural-language predicate. Both can use a cheap proxy first and a stronger model second. LOTUS’s
sem_filterprototype isl_M(t_i)over each tuple; ourwhere decideis that filter with a decision as the predicate. - Narrower. LOTUS’s
sem_filterreturns a Boolean per record. Our filter’s predicate is a typed decision: it can carry a calibrated probability and an abstention, and the abstention is exercised (the removed records are visible asabstained, not silently dropped). A Boolean cannot represent “I am not sure”. - Different. LOTUS specifies a gold algorithm per operator and provides statistical accuracy guarantees (recall and precision targets
gamma_R, gamma_Pmet with probability1 - delta) via sampling on the proxy, with a cost-based optimizer over an operator algebra. Our decision filter uses the Chapter 11 risk machinery instead of LOTUS’s sampling framework, and lives in a language, not a data-processing engine. LOTUS’s gold algorithm is an LLM call per tuple; ours is a CPU cross-encoder.
We claim no novelty. where decide is sem_filter with a typed, abstaining predicate and a different guarantee machinery. The chapter’s result is not that the construct is new; it is that, on this task, the filter choice matters more than the syntax, and that the risk-controlled predicate is the wrong operating point.
Wrong / Correct. Wrong: “
where decideis a new semantic operator; the syntax is the contribution.” Correct:where decideis LOTUS’ssem_filterunder another name. The contribution, if any, is the typed abstaining predicate and the calibration discipline — and on this task that discipline hurt.
The distinction this chapter keeps
A cross-encoder is a relevance scorer, not a decision model. It produces a ranking score. Calling its thresholded output a “decision” is a use of the contract, not evidence about a decision model. The typed decision (probability, abstention, refusal) is what the contract adds; the scorer is what the provider is.
Filter quality vs filter cost. The cascade is the Pareto point here: cross-encoder quality at half the passes. The risk-controlled filter is a different point (precision-first) and is not comparable on end-to-end accuracy alone.
What to carry forward
where decide is a method, and the library form is sem_filter plus a threshold. The construct earns syntax only if it enforces a discipline the library form cannot, and that discipline measurably prevents failure. The abstention discipline is the candidate; on this evidence it did not prevent failure, it caused it. That is a result for the language hypotheses (L1) and it is decided (if at all) in Chapter 35.
Chapter 23 asks the next question: if filtering is a method, what about iteration — is for decide a construct or sugar for filter-then-loop?
Close by
Does it deserve syntax, or is it a method? On this evidence, a method. The next chapter tests the same question for iteration.
Limitations
- One dataset (SciFact), one encoder, one cross-encoder, one CPU machine. Not a claim about filters in general.
- The decision filter’s conservatism is partly a configuration choice (target risk 0.20, delta 0.10). A higher risk target would trade precision for coverage; the chapter reports the configured point only.
- The cross-encoder is used zero-shot, not fine-tuned on SciFact; a fine-tuned cross-encoder could change the ranking.
- The cascade prefilter size m=5 is fixed, not tuned; the cascade’s cost saving is reported for that m.
- LOTUS read in part (Abstract, §1, §2.1-§2.3, §3.1). Nogueira and ColBERT were read in full text after drafting (Session R); LOTUS §5 (evaluation) remains unread, and no claim here rests on it.
where decideis not implemented as syntax; the chapter tests the filter, and infers the construct question from it.