← Jev From First Principles

Semantic Pipelines

Can a real task be built from retrieval, relevance, relation, generation and support decisions?

Semantic Pipelines

This chapter reports a CPU run of the pipeline’s retrieval and relevance stages over SciFact, and the full pipeline with a declared weak relation provider. The relation (entailment) and generation providers need a language model and are deferred. The author’s private Writer evidence is off limits; the public claims are SciFact’s.

The problem

Chapters 21-24 built the pieces: candidate generation, filtering, relation, composition. This chapter asks whether they fit together into a real task. The task is claim research: given a claim, retrieve evidence, decide whether the evidence supports or contradicts the claim, produce an answer, and decide whether that answer is supported.

    flowchart LR
    claim --> R[retrieve]
    R --> F[relevance]
    F --> Rel[relation]
    Rel --> G[generate]
    G --> S[support]
    S --> verdict[verdict]
  

The chapter question: can a real task be built from these decisions, and is the composition auditable?

What we expect and why

Our starting hypothesis is that composition is cheaper and auditable, a single large call is competitive on accuracy, and errors concentrate in the relation decision.

Three papers frame it:

  1. Thorne et al. (2018) — FEVER is the canonical fact-verification task: claims labelled Supported, Refuted or NotEnoughInfo against Wikipedia evidence, with the evidence sentences recorded. The task and its data. (Read at abstract level.)

  2. Gao et al. (2022) — RARR researches and revises generated text against evidence, post-editing unsupported content while preserving the original. An attribution pipeline is the prior art shape. (Read at abstract level.)

  3. Asai et al. (2023) — Self-RAG shows a single LM can learn to retrieve and critique using reflection tokens. This challenges the composed decision view: the critique can be inside one model rather than spread across stages. (Read at abstract level.)

The first gives the task, the second the shape, the third the counterposition.

The build

src/arbiter/pipeline.py is typed plumbing: five stages, each a transformation over a ClaimState plus a trace entry. The providers are callables, so the same pipeline runs with a real retriever or with stubs. The point is auditability: attribute_error names the one stage that caused an end-to-end error.

from arbiter.pipeline import attribute_error, run_pipeline

DOCS = ["the cat sat", "the dog sat", "a cat and a dog"]
    # 1. Build the pipeline from supplied providers (no model).
    print("1. run the pipeline over tiny docs (hand-supplied providers)")

    def retriever(claim):
        return [0, 1, 2]

    def filter_fn(claim, candidates):
        return [0, 2]  # kept the cat docs

    def relation_fn(claim, kept_idx):
        return "SUPPORT"  # a supplied relation decision

    def generator(state):
        return "the claim is " + state.relation.lower() + "ed"

    def checker(state):
        return True

    run = run_pipeline("the cat", retriever, filter_fn, relation_fn, generator, checker)
    for t in run.trace:
        print(f"   {t.stage:<9} -> {t.detail}")
    print(f"   verdict = {run.verdict}")
    # 2. Error attribution: which stage caused an error?
    print("2. error attribution")
    state = run.state
    for gold_ev, gold_rel, label in (
        ({9}, "SUPPORT", "evidence 9 never retrieved"),       # -> retrieve
        ({1}, "SUPPORT", "evidence 1 in candidates, not kept"),  # -> relevance
        ({0}, "CONTRADICT", "evidence 0 kept, relation disagrees"),  # -> relation
    ):
        stage = attribute_error(gold_ev, state, gold_rel)
        print(f"   {label:<42} -> {stage}")

The walkthrough prints:

1. run the pipeline over tiny docs (hand-supplied providers)
   retrieve  -> 3 candidates
   relevance -> 2 kept
   relation  -> SUPPORT
   generate  -> the claim is supported
   support   -> True
   verdict = SUPPORT
2. error attribution
   evidence 9 never retrieved                 -> retrieve
   evidence 1 in candidates, not kept         -> relevance
   evidence 0 kept, relation disagrees        -> relation

The experiment

python examples/ch25-semantic-pipelines/run_ch25.py runs the pipeline on the 188 retrievable SciFact dev claims. Retrieval is BM25 top-10; relevance is the cached ms-marco cross-encoder at the Chapter 22 threshold; the relation stage is a declared weak lexical baseline (Jaccard overlap of content words at 0.3), not an NLI model; generate is a stub template; support is a typed check. The NLI and generation providers are PENDING_RUN.

Test split (results/ch25.jsonl):

Quantity Value
end-to-end verdict accuracy 0.165
per-stage error rate: retrieve 0.080
per-stage error rate: relevance 0.314
per-stage error rate: relation 0.441
relation stage accuracy (weak baseline) 0.197

What surprised us

  1. Every error is attributable to exactly one stage. No claim had an ambiguous failure. The composition is auditable, which is the strongest thing this chapter shows: the partition {retrieve, relevance, relation} covers all 157 errors, 31 claims correct, 188 total.

  2. Errors concentrate in relation, but partly because the supplied baseline is weak. The lexical relation baseline is worse than predicting the majority label (0.197 vs the 0.64 majority base rate on SciFact dev). The concentration of errors in relation is real, but it is a property of the weak provider we were allowed to run, not a claim about relation decisions in general.

  3. Relevance is the second-largest error source, and it is a real CPU measurement. 0.314 of claims had their evidence retrieved but dropped by the cross-encoder filter — the Chapter 22 decision-error story, now as a pipeline number.

  4. Retrieval is the smallest error source. 0.080 of claims lacked the evidence in the top-10. This is consistent with Chapters 21-22: the retriever is not the bottleneck.

  5. The single-call baseline is not runnable here. The prompt asks for a single LLM call and a Self-RAG-style one-model critique; both need a language model and are NOT_OBSERVED. Saying that is the honest result, not a gap to paper over.

Wrong / Correct. Wrong: “The pipeline beats a single call because composition is better.” Correct: “The pipeline is auditable — every error has a stage — and the single call is competitive or better on accuracy. What composition buys is attribution, not accuracy. Neither claim is measured here beyond the attribution.”

The distinction this chapter keeps

A pipeline is software; a pipeline run is evidence. The typed plumbing is tested and real. The per-stage attribution is a real CPU measurement. The 0.165 end-to-end number, however, is the pipeline with a weak supplied relation provider; it is not a claim about what a real relation decision achieves.

Retrieval error, relevance error and relation error are separate budgets. They have different sizes (0.080, 0.314, 0.441) and different fixes: a bigger candidate list, a better filter, an entailment model.

What to carry forward

Which parts of the pipeline justified being decisions? On this evidence, the retrieval and relevance stages justify being stages (they are real, measurable, and their errors are separable). The relation stage is where a decision model would earn its place, and it is the one we could not run. The generation stage is a stub. Chapter 26 generalises the pipeline from a list to a graph, where the trace this chapter built becomes the graph’s provenance.

Close by

Which parts of the pipeline justified being decisions? All of them, as stages — and precisely the relation stage is the one that needs a real provider. Chapter 26 builds the graph that carries every stage’s provenance forward.

Limitations

  • Retrieval and relevance are real CPU measurements; relation is a declared weak lexical baseline; generate is a stub. The NLI and generation providers are PENDING_RUN.
  • The single-call and Self-RAG-style baselines are NOT_OBSERVED (need an LLM).
  • The author’s private Writer evidence is off limits (hand-off); only public SciFact claims were used.
  • The three required papers were read in full text after drafting (Session R), which found no contradiction with the chapter’s claims.
  • The relation error concentration (0.441) is partly an artifact of the weak baseline; a real entailment provider would change it.