for decide
Does semantic iteration deserve its own syntax, or is it sugar for filtering?
for decide
This chapter reports a CPU experiment that ran here: the loop strategies over the Chapter 22 SciFact candidate lists and their cached cross-encoder scores. No language model was run in this chapter. The result is a rejection of the construct, which the chapter is allowed to make.
The problem
Chapters 21 and 22 showed that something must select what a decision reads, and that the filter is where the end-to-end loss lives on SciFact. This chapter asks the next question: is iteration its own construct?
for decide relevant(doc, claim) in corpus:
act(doc)
against its library form:
for doc in filter(relevant, corpus):
act(doc)
The chapter question is whether the construct expresses anything the library cannot.
What we expect and why
Our starting hypothesis is that only early termination and ranked iteration differ, and both are better served by library calls.
Three papers frame the comparison:
-
Shankar et al. (2024) — DocETL optimises complex document-processing pipelines and finds plans 25–80% more accurate than well-engineered baselines, using agent-based rewrites that decompose the data or the task. The prior art is optimisation of semantic operators, not a new iteration construct. (Read at abstract level.)
-
Sun et al. (2023) — RankGPT re-ranks a candidate set with an LLM in one pass and reports competitive or superior results to supervised rankers; a distilled 440M model beats a 3B supervised model on BEIR. Ranking is a decision over a set. (Read at abstract level.)
-
Qin et al. (2023) — Pairwise Ranking Prompting argues that listwise prompts are hard for LLMs and pairwise comparisons can beat listwise ranking, with linear-complexity variants. This contradicts RankGPT’s listwise framing: the two papers disagree on how to order a set. (Read at abstract level.)
Both rankers order a candidate list before iterating it. That is the shape the construct claims: an iteration whose order is itself decided.
The build
src/arbiter/iteration.py implements five loop strategies and a comparator.
from arbiter.iteration import (
filter_then_loop,
for_decide,
for_decide_early,
for_decide_ranked,
library_early_stop,
ranked_is_sort_plus_loop,
)
DOCS = ["no", "yes", "no", "yes", "no"]
is_yes = lambda d: d == "yes"
# 1. The two forms agree when neither stops early.
print("1. eager filter-then-loop vs lazy for-decide")
eager = filter_then_loop(DOCS, is_yes)
lazy = for_decide(DOCS, is_yes)
print(f" decision equal: {eager.decision == lazy.decision}")
print(f" evidence equal: {eager.evidence == lazy.evidence}")
print(f" docs consulted: eager={eager.docs_consulted} lazy={lazy.docs_consulted}")
# 2. Early termination consults fewer documents.
print("2. early termination (stop at the first relevant doc)")
early = for_decide_early(DOCS, is_yes)
print(f" decision={early.decision} evidence={early.evidence} consulted={early.docs_consulted}")
# 3. The library form is identical.
print("3. the library form (takewhile)")
lib = library_early_stop(DOCS, is_yes)
print(f" consulted equal: {early.consulted == lib.consulted}")
print(f" evidence equal: {early.evidence == lib.evidence}")
# 4. Ranked iteration is a sort plus the same loop.
print("4. ranked iteration")
scores = [0.1, 0.5, 0.2, 0.9, 0.0]
ranked = for_decide_ranked(DOCS, scores, is_yes)
print(f" visit order (original indices): {ranked.consulted}")
print(f" is sort+loop: {ranked_is_sort_plus_loop(DOCS, scores, is_yes)}")
# 5. The decision the construct makes, and the library form of it.
print("5. the decision, and its library form")
print(f" for decide relevant(doc) in corpus -> {early.decision}")
print(f" any(relevant(d) for d in corpus) -> {any(is_yes(d) for d in DOCS)}")
The walkthrough prints:
1. eager filter-then-loop vs lazy for-decide
decision equal: True
evidence equal: True
docs consulted: eager=5 lazy=5
2. early termination (stop at the first relevant doc)
decision=True evidence=('yes',) consulted=2
3. the library form (takewhile)
consulted equal: True
evidence equal: True
4. ranked iteration
visit order (original indices): (3,)
is sort+loop: True
5. the decision, and its library form
for decide relevant(doc) in corpus -> True
any(relevant(d) for d in corpus) -> True
The rejection rule, stated before the comparison
REJECT for decide if: (a) its decision and evidence equal filter-then-loop’s whenever neither stops early, and (b) every difference is reproduced by a library call — early termination by takewhile/any, ranked order by sorted — at equal or lower cost. ADOPT if some task needs a behaviour neither the library form nor a sort provides.
The experiment
python examples/ch23-for-decide/run_ch23.py runs the strategies over the Chapter 22 candidate lists: SciFact, BM25 top-10 per claim, the same 188 test claims, with the cached ms-marco cross-encoder score per candidate and the Chapter 22 threshold (tau 3.3733) deciding “evidence”.
Test split (results/ch23.jsonl):
| Quantity | Value |
|---|---|
| decision agreement (all four strategies) | 1.000 |
| evidence equality, eager vs lazy | 1.000 |
| ranked iteration equals sort + loop | 1.000 |
| mean documents consulted (eager) | 10.00 |
| mean documents consulted (lazy, no stop) | 10.00 |
| mean documents consulted (early stop) | 4.04 |
| mean documents consulted (ranked) | 3.92 |
mean documents consulted (library takewhile) |
4.04 |
Applying the rule: (a) holds (1.000 evidence equality and decision agreement); (b) holds (the library form consults exactly 4.04 documents, the same as the construct’s early stop; the ranked form is a sort plus the same loop, and the comparator returns true on every claim). The verdict is REJECTED.
What surprised us
-
Lazy and eager are observationally identical with a pure predicate. The construct’s supposed advantage over
filteris that it interleaves evaluation with action. With a pure predicate that interleaving has no observable effect: same decision, same evidence, same count (10.00 both). The difference only appears with early termination — which is a separate mechanism, not iteration order. -
Early termination is a library call.
takewhileconsults exactly as many documents as the construct’s early stop (4.04 both), on all 188 claims. Nothing about the language form is needed to get the saving. -
Ranked iteration barely helps, and it is a sort. Ranking the candidates by the cross-encoder before iterating saved only 0.12 documents on average (4.04 to 3.92), because the cross-encoder predicate itself misses evidence for about a third of claims (the Chapter 22 decision error), and those claims scan the whole list either way. The ranked form is exactly
sorted(...)followed by the same loop (ranked_is_sort_plus_loopreturns true on every claim), so it adds no semantics beyond a sort. -
The decision it computes is
any. With the default sufficiency (one relevant document is enough), the construct’s decision isany(relevant(d) for d in corpus)— a library call, with no new type and no new failure mode.
Wrong / Correct. Wrong: “
for decideis a construct because it lets the program act while it searches.” Correct: on a pure predicate, acting while searching isfilterplustakewhile. The construct’s only behaviour the library lacks is early termination, and that isany/takewhile. The verdict is REJECTED.
Prior art
DocETL optimises pipelines by rewriting them (decomposing the data or the task); it does not introduce a new iteration construct. RankGPT and PRP order a candidate set before using it — a ranking step, not a loop primitive — and they disagree with each other on whether the ordering decision should be listwise or pairwise. Nothing in the three papers requires a new loop form, and this chapter claims none.
The distinction this chapter keeps
A loop over decisions is not a decision. for decide mixes iteration (a control-flow concern) with the semantic decision (a provider concern). The decision belongs to the predicate; the iteration belongs to for. Keeping them together in one keyword is what makes the construct look like it adds something, when the two halves are separately expressible.
Cost is not semantics. Early termination changes how many documents are consulted; it does not change the answer. A construct that only changes cost is an optimisation, and optimisations belong in the library or the compiler, not the syntax.
What to carry forward
If for decide is rejected, which constructs remain? On this evidence, decide, if decide and where decide stand (Chapters 16, 17, 22), and for decide becomes for ... in filter(...) plus takewhile. The next chapter asks a different question: how do decisions about decisions compose, and does a chain of decisions need a controller?
Close by
If this is rejected, which constructs remain? The iteration form does; the language construct does not. Chapter 24 turns to composition.
Limitations
- One corpus (SciFact), one scorer (ms-marco cross-encoder), one threshold (tau 3.3733 from Chapter 22). The loop comparison is deterministic given the scores.
- RankGPT and PRP were read in full text after drafting (Session R); DocETL is PARTIAL (Abstract, §1-§3 read; §4-§5 unread) and the chapter’s DocETL claims rest on §1-§3. Session R found no contradiction with the chapter.
- The early-termination saving depends on the scorer’s decision error: with a perfect predicate the first document would nearly always suffice. The 4.04 figure is a property of this scorer and corpus.
- The rejection rule was written before the comparison and is applied as written; a different rule (for example, adopting on any cost saving) would give a different verdict.