Embeddings as Candidate Generation
How do retrieval errors and decision errors combine when a decision can only read a candidate set?
Embeddings as Candidate Generation
This chapter reports a CPU experiment that ran here: an open encoder (
bge-small-en-v1.5) and BM25 over a small public dataset (SciFact), on the machine described inevidence/environment.md. It also shows a model-free walkthrough. The encoder is a relevance model, not a decision model: nothing here is evidence about Jev or about H1-H4.
The problem
Chapter 20 ended with a claim: a decision cannot read everything. If the state is larger than the model can use, the decision layer must be fed a smaller state. Something has to choose what it reads.
This chapter measures the two places the pipeline can fail:
- the candidate generator does not retrieve the evidence at all (retrieval error), and
- the decision layer drops evidence that was retrieved, or keeps distractors (decision error).
The chapter question is exactly this: where is the end-to-end error — in what we retrieve, or in what we decide?
What we expect and why
Our starting hypothesis is that most end-to-end errors are retrieval misses for rare evidence, and that a hybrid generator beats either retriever alone.
Three papers frame the expectation:
-
Karpukhin et al. (2020) — Dense Passage Retrieval (DPR) uses two BERT encoders whose dot product is a decomposable similarity; it beats a tuned Lucene-BM25 by 9-19% top-20 on five QA datasets. The same paper shows that a linear combination of BM25 and dense scores (“BM25+DPR”) sometimes beats either alone — the origin of the hybrid expectation.
-
Thakur et al. (2021) — BEIR evaluates ten retrieval systems over 18 heterogeneous datasets and reports that “BM25 is a robust baseline”, while “dense and sparse-retrieval models … often underperform”. Dense retrieval does not dominate everywhere. This is the source that weakens our expectation: the candidate generator may not be where quality is won.
-
Lewis et al. (2020) — RAG fixes the pipeline shape: a retriever selects passages, and a generator conditions on them. We keep the shape and replace the generator with a decision. (Read at abstract level; cited only for that shape.)
The shape of the pipeline
flowchart LR
Q[claim] --> G[candidate generator]
C[corpus] --> G
G --> S[candidate state, top-k]
S --> D[decision layer]
D --> O[verdict]
The generator is BM25, dense, or hybrid. The decision layer is the simplest relevance decision: keep a candidate if an embedding relevance score clears a threshold. Chapter 22 compares better filters.
The build
src/arbiter/retrieve.py owns the candidate generator and the retrieval metrics. BM25 is written in the repo (inverted index, tested against a hand-computed example) rather than pulled from a package. The dense index takes an encoder passed in by the caller, so the module never downloads anything.
The walkthrough is model-free: a tiny corpus, hand-checkable scores. It shows the BM25 arithmetic, the ranking rule, the metrics, the fusion, and the 2x2 breakdown.
from arbiter.retrieve import (
BM25Index,
hit_at_k,
mrr,
ndcg_at_10,
paired_bootstrap,
rank,
recall_at_k,
rrf_fuse,
)
DOCS = ["the cat sat", "the dog sat", "a cat and a dog"]
# 1. BM25 over a tiny corpus, checkable by hand.
print("1. BM25 over a tiny corpus")
idx = BM25Index(k1=0.9, b=0.4).fit(DOCS)
print(" IDF('cat') =", round(idx._idf["cat"], 6), "(= ln(1.6))")
scored = idx.search("cat", 3)
for i, s in scored:
print(f" doc {i} ({DOCS[i]!r}): {s:.5f}")
# 2. rank breaks ties toward the lower index.
print("2. rank")
print(" rank([0.5, 0.5, 0.1], 2) =", rank([0.5, 0.5, 0.1], 2))
# 3. Retrieval metrics on one query with two relevant docs.
print("3. metrics (relevant = {0, 2})")
ranked = idx.search("cat", 3)
rel = {0, 2}
for k in (1, 3):
print(f" k={k}: recall={recall_at_k(ranked, rel, k):.2f} hit={hit_at_k(ranked, rel, k):.2f}")
print(f" MRR={mrr(ranked, rel):.3f} nDCG@10={ndcg_at_10(ranked, rel):.3f}")
# 4. Hybrid fusion uses ranks, not raw scores.
print("4. RRF fusion (dense weight w)")
dense = [0.9, 0.1, 0.5]
sparse = [0.0, 5.0, 1.0]
for w in (0.0, 0.5, 1.0):
top = rank(rrf_fuse(dense, sparse, weight=w), 1)[0][0]
print(f" w={w}: top doc = {top}")
# 5. The 2x2 breakdown: retrieval (found/missed) x decision (kept/dropped).
print("5. 2x2 error breakdown (hand-supplied)")
cases = {
"retrieved+kept": (True, True),
"retrieved+dropped": (True, False),
"missed": (False, None),
}
counts = {"retrieved_kept": 0, "retrieved_dropped": 0, "missed": 0}
for name, (retrieved, kept) in cases.items():
if not retrieved:
counts["missed"] += 1
elif kept:
counts["retrieved_kept"] += 1
else:
counts["retrieved_dropped"] += 1
n = len(cases)
print(f" retrieved+kept = {counts['retrieved_kept']} (end-to-end correct)")
print(f" retrieved+dropped = {counts['retrieved_dropped']} (decision error)")
print(f" missed = {counts['missed']} (retrieval error)")
print(f" end-to-end = {counts['retrieved_kept'] / n:.3f}")
# 6. Paired bootstrap on a difference.
print("6. paired bootstrap")
a = [1.0, 1.0, 0.0, 1.0]
b = [0.0, 1.0, 0.0, 1.0]
diff, lo, hi = paired_bootstrap(a, b, n_boot=2000, seed=0)
print(f" mean diff = {diff:+.3f} 95% CI [{lo:+.3f}, {hi:+.3f}]")
The walkthrough prints:
1. BM25 over a tiny corpus
IDF('cat') = 0.470004 (= ln(1.6))
doc 0 ('the cat sat'): 0.48677
doc 2 ('a cat and a dog'): 0.43971
doc 1 ('the dog sat'): 0.00000
2. rank
rank([0.5, 0.5, 0.1], 2) = [(0, 0.5), (1, 0.5)]
3. metrics (relevant = {0, 2})
k=1: recall=0.50 hit=1.00
k=3: recall=1.00 hit=1.00
MRR=1.000 nDCG@10=1.000
4. RRF fusion (dense weight w)
w=0.0: top doc = 1
w=0.5: top doc = 0
w=1.0: top doc = 0
5. 2x2 error breakdown (hand-supplied)
retrieved+kept = 1 (end-to-end correct)
retrieved+dropped = 1 (decision error)
missed = 1 (retrieval error)
end-to-end = 0.333
6. paired bootstrap
mean diff = +0.250 95% CI [+0.000, +0.750]
The experiment
The run is python examples/ch21-embeddings-as-candidate-generation/run_ch21.py. It is CPU-only, resumable (corpus embeddings cache to results/partial/ch21/), and writes results/ch21.jsonl.
- Corpus and queries. SciFact: 5,183 abstracts; claims with a SUPPORT or CONTRADICT evidence document are the retrievable queries. Train pool 505 retrievable claims, shuffled at seed 0 and cut into
train105 /calibrate200 /threshold200.testis the 188 retrievable SciFact dev claims, touched once. - Generators. BM25 (grid over k1, b); dense (
bge-small-en-v1.5, revision5c38ec7c..., no hyperparameters); hybrid (reciprocal-rank fusion, weight w). The BM25 grid and the fusion weight are both declared and selected oncalibrateby the same rule (hit@10). The dense generator isUNTTUNED, as it has no hyperparameters to tune and training it is out of scope. - k. Chosen by a declared rule: the smallest k with
calibratehit@k at least 0.90. For BM25 that is k=10; for dense and hybrid, k=5. - Decision. A per-pair relevance filter: keep a candidate if the dense cosine between claim and document is at least tau. tau is selected on
thresholdto maximise pair-level F1. Because the candidate set contains non-evidence documents, the filter has real negatives and tau is not degenerate. Selected tau=0.81 (pair-F1 0.608, precision 0.568, recall 0.652).
The result
Retrieval on the test split (results/ch21.jsonl, stage=retrieval):
| Generator | recall@10 | hit@10 | MRR | nDCG@10 |
|---|---|---|---|---|
| BM25 (k1=1.2, b=0.4) | 0.905 | 0.920 | 0.792 | 0.813 |
| dense (bge-small) | 0.953 | 0.963 | 0.835 | 0.859 |
| hybrid (w=0.75) | 0.961 | 0.973 | 0.832 | 0.859 |
The hybrid beats either generator alone on hit@10. The paired bootstrap against BM25 puts the dense gain at +0.043 [0.000, 0.085] — the interval touches zero — and the hybrid gain at +0.053 [0.021, 0.090], which excludes zero (stage=retrieval_paired_vs_bm25). At n=188, the dense advantage is not established; the hybrid advantage is small but clear.
The 2x2 error breakdown (stage=error_2x2) is where the chapter’s question is answered:
| Generator | end-to-end kept | decision error | retrieval error | decision P / R |
|---|---|---|---|---|
| BM25 | 0.559 | 0.362 | 0.080 | 0.567 / 0.526 |
| dense | 0.569 | 0.335 | 0.096 | 0.542 / 0.560 |
| hybrid | 0.569 | 0.367 | 0.064 | 0.553 / 0.545 |
What surprised us
-
Decision error dwarfs retrieval error. For every generator the decision layer drops 0.34-0.37 of the test claims whose evidence was retrieved, while retrieval misses only 0.06-0.10. Our preregistered expectation — that most end-to-end errors are retrieval misses — is refuted on this task. The weak embedding filter is the bottleneck, not the retriever. This is the opposite of the intuitive reading and it is what the numbers say.
-
The hybrid helps, but only slightly. The hybrid is the best generator (hit@10 0.973), and its paired interval excludes zero. But the dense generator alone is within its own interval of the hybrid, consistent with BEIR’s warning that dense does not dominate and that the differences are task-dependent.
-
BM25 is competitive. BM25 hit@10 is 0.920 against dense’s 0.963. On a corpus of scientific abstracts with distinctive terminology, a lexical baseline that has been strong for decades is close to a modern encoder. That matches BEIR and matches DPR’s own SQuAD caveat.
-
The relevance decision is weak. At the F1-optimal threshold the filter keeps roughly half the evidence pairs (recall ~0.55) and about half of what it keeps is non-evidence (precision ~0.55). A thresholded embedding cosine is not a good relevance decision. That is the finding Chapter 22 chases.
Wrong / Correct. Wrong: “retrieval is the bottleneck; improve the retriever and the pipeline improves.” Correct: on this task the retrieval stage retains 92-97% of the evidence, and the decision stage throws away a third of what retrieval found. The bottleneck is the decision.
The distinction this chapter keeps
Retrieval error is not decision error. They are counted separately at the query level: a claim whose evidence was never retrieved is a retrieval error; a claim whose evidence was retrieved and then dropped is a decision error. They have different sizes here (0.06-0.10 vs 0.34-0.37) and different fixes.
Candidate generation is not a decision. DPR-style encoders and BM25 rank; they do not decide relevance with a calibrated, thresholded, abstaining answer. Nothing measured here is evidence about any decision model; the cross-encoder in Chapter 22 is still a relevance scorer.
What to carry forward
If retrieval caps the decision, the decision layer’s real contribution is to filter what retrieval over-recalls — and the filter matters more than the generator on this task. Chapter 22 asks whether a better filter (a cross-encoder, a decision filter, or a cascade) actually reduces the decision error, and whether where decide is a construct or just a method.
Close by
If recall caps the decision, what is the decision layer’s real contribution? On this evidence, the decision layer’s contribution is negative — it removes more evidence than the retriever misses. The next chapter tests whether a better filter can turn that around.
Limitations
- One dataset (SciFact, scientific abstracts), one encoder (bge-small), one CPU machine. Not a claim about retrieval in general.
- The decision layer is the embedding filter only. A cross-encoder or a trained filter could reverse the decision-error finding; Chapter 22 tests that.
- The SciFact upstream test split has no public labels, so
testhere is the labelled dev split (188 retrievable claims). The runner’s first execution preceded the preregistration commit, and a run evaluates every split, including the labelled dev split used astest. That split was therefore evaluated more than once: in the first run and again in the post-preregistration re-run whose rows are reported. The run is deterministic (seed 0, cached embeddings), so the re-run reproduces the first, but the ordering was wrong and the “once” in the protocol was not kept. The decision step was also first defined in a way that was degenerate on the threshold split (every claim there had evidence, so a presence decision was trivial) and was redefined as a per-pair filter with real negatives before any number was reported. - tau is chosen by pair-F1 on
threshold, which balances precision and recall. A different target (e.g. high precision) would change the decision error. - The BM25 grid and hybrid weight are selected on
calibrate; dense is untuned. The comparison is symmetric on the tuning rule, not on the number of configurations. - The three required papers: DPR, BEIR and RAG were read in full text (BEIR §1, §3-§6; RAG §1-§4). The chapter’s claims about BEIR and RAG were written from abstracts and checked against the full text in Session R, which found no contradiction.