← Jev From First Principles

The Decision Compiler

Can a compiler pick the cheapest provider that satisfies a decision contract?

The Decision Compiler

This chapter is a simulation over supplied profiles. The provider curves, costs and the cascade’s end-to-end profile are seeded from earlier rows (Ch04’s measured LR; Ch30’s simulation rows) or are explicit ASSUMPTIONS. The compile-time table is arithmetic on those profiles; the runtime numbers are a SIMULATION over the same assumed curves. No new model runs.

The problem

A decision contract is now declarative: Chapter 32’s DecisionSpec names the inputs, answers, risk, abstention route and fallback chain. This chapter asks the next question: given the spec and a set of provider profiles, can a compiler pick the cheapest provider that satisfies the contract?

The prompt’s own question is a good guardrail: the compiler may conclude that it adds nothing over a good default cascade. That conclusion would be a result.

What we expect and why

Our starting hypothesis is that a cost-based compiler matches an expert’s hand-built plan when profiles are accurate and is worse when they drift, and that the two decision times must be kept apart: provider choice happens at compile time from profiles; escalation on uncertainty happens at run time.

Three sources frame the mechanism:

  1. Selinger et al. (1979) — System R chooses access paths by estimated cost from a declarative query. The model for compile-time selection. This is a classic; it was verified at DOI level (10.1145/582095.582099) before citation, not quoted from memory.
  2. Liu et al. (2024), Palimpzest (2405.14696) — a declarative language over semantic operators with a cost-optimization framework that enumerates plans trading runtime, financial cost and output quality; up to 3.3x faster and 2.9x cheaper than baseline on its tasks. The closest prior art for this chapter: same ideal, larger scope. (Read at abstract level.)
  3. Liu et al. (2024), relational analytics (2403.05821) — concrete serving-level cost reduction by reordering rows and fields to reuse KV caches (up to 3.4x faster, 32% cheaper). A different mechanism from ours, and the boundary is stated: we are not comparing serving-level optimizations. (Read at abstract level.)

The paper that weakens the expectation is Palimpzest itself: its speedups come from planning across many operators at workload scale, so a single-decision contract may simply have no cascade worth choosing — the “compiler adds nothing over a default” branch is live, not decorative. We check it.

The build

src/arbiter/compiler.py implements compile(spec, profiles) -> plan. Selection rule: expected accuracy over uniform difficulty is (p_hi + p_lo)/2 for singles; the cascade candidate uses its supplied end-to-end accuracy — never the naive product of stage accuracies, because Chapter 30 measured that product does not compose (0.577 naive vs 0.830 measured). The cheapest feasible candidate wins; if none is feasible, the plan escalates to “human always”. The plan separates compile-time choices from runtime rules.

from arbiter.compiler import CascadeProfile, ProviderProfile, compile_plan, explain
from arbiter.spec import AnswerClass, DecisionSpec
    # 1. compile(spec, profiles): cheapest feasible single
    print("1. compile(spec, profiles)")
    plan = compile_plan(spec_for(0.85), profiles(), cascade())
    print(f"   chosen {plan.steps[0].provider}: expected accuracy "
          f"{plan.expected_accuracy:.3f} >= {plan.quality_target:.3f}, "
          f"cost {plan.expected_cost:.2f}")
    print(f"   cascade expected {cascade().accuracy:.3f} < {plan.quality_target:.3f}"
          f" -> infeasible (never the naive product)")
    # 2. drift 0.9 on every profile flips the choice
    print("2. drift 0.9 on every profile")
    drifted = compile_plan(spec_for(0.85), profiles(), cascade(), multiplier=0.9)
    print(f"   chosen {drifted.steps[0].provider}: expected accuracy "
          f"{drifted.expected_accuracy:.3f} (linear marginal 0.774 < 0.85), "
          f"cost {drifted.expected_cost:.2f}")
    # 3. explain(plan) states the evidence behind every choice
    print("3. explain(plan)")
    for line in explain(plan):
        print("   " + line)
1. compile(spec, profiles)
   chosen linear: expected accuracy 0.860 >= 0.850, cost 5.00
   cascade expected 0.830 < 0.850 -> infeasible (never the naive product)
2. drift 0.9 on every profile
   chosen judge: expected accuracy 0.860 (linear marginal 0.774 < 0.85), cost 20.00
3. explain(plan)
   plan for 'policy' (quality target 0.850)
   expected accuracy 0.860, expected cost 5.00 per request
   compile-time choices:
     - linear chosen: expected accuracy 0.860 >= target 0.850; cost 5.00; cheapest feasible: expected accuracy 0.860 >= target 0.850
   runtime rules:
     - none (single provider, always answers)
   profile evidence:
     - rule: assumption
     - linear: Ch04 LR 0.8779 seed
     - judge: assumption
     - human: assumption
     - cascade: seeded from results/ch30.jsonl (mode=simulation, stage=cascade, answer_fraction 0.5, drift 1.0)

The run

run_ch33.py writes results/ch33.jsonl (mode analytic + mode simulation).

Compile-time table (arithmetic on the supplied profiles):

target chosen expected accuracy expected cost
0.80 linear 0.860 5.00
0.85 linear 0.860 5.00
0.90 judge 0.955 20.00
0.95 judge 0.955 20.00

The cascade (supplied end-to-end 0.830, cost 11.797) is never chosen at any target in this table: at every target where it is feasible, a single provider is feasible at or below its cost, and where it is infeasible nothing changes. Given these profiles, the hand-built cascade of Chapter 30 is Pareto-dominated — cheaper and less accurate than the single linear provider. That is a property of the supplied profiles, stated as such.

Drift recompile (multiplier on every quality profile, target 0.85):

drift chosen expected accuracy expected cost
0.9 judge 0.860 20.00
1.0 linear 0.860 5.00
1.1 linear 0.913 5.00

A 0.9 quality drift drops linear’s marginal to 0.774, below the contract, so the compiler recompiles to judge — accuracy held at the same expected value, cost quadrupled.

Realized (SIMULATION, n=4000, the contracts’ own target 0.85):

world compiled plan cheapest-adequate (frozen) always-strongest hand-built cascade (frozen)
drift 0.9 judge 0.859 · 20.00 · held 0.778 · 5.00 · broken 0.859 · 20.00 · held 0.797 · 22.90 · broken
drift 1.0 linear 0.861 · 5.00 · held 0.861 · 5.00 · held 0.959 · 20.00 · held 0.830 · 11.80 · broken
drift 1.1 linear 0.913 · 5.00 · held 0.913 · 5.00 · held 1.000 · 20.00 · held 0.859 · 5.97 · held

(accuracy · cost per request · contract held or broken at 0.85)

What it says

  1. The compiler picks the cheapest provider that satisfies the contract — and that is not the expert cascade. Given the Chapter 30 profiles, the compiled plan is the single linear classifier at target 0.85: realized 0.861 accuracy, 5.00 per request. The hand-built cascade runs 0.830 at 11.80 — it misses the 0.85 contract on the profiles it was built from (Chapter 30 already showed its end-to-end guarantee does not compose: 0.577 naive vs 0.830 measured). Cost-based selection, left to the arithmetic, reproduces what a human would likely not write down: the cascade is dominated. P1, P2 hold.

  2. The cascade is never optimal under these profiles. (Read this narrowly: the linear profile is a measured Chapter 4 number on its own task, while the cascade’s 0.830 comes from Chapter 30’s assumed stage curves. The two were never run on the same data, so this compares a measurement with an assumption, and says nothing about real cascades.) P1’s “cascade is chosen only when no single is feasible and it clears the target” branch is unreachable here: the cascade loses to linear on cost and accuracy simultaneously. The chapter that asked “did a good default cascade exist” gets the honest answer: under the supplied profiles, the good default is a single adequate provider, and the compiler finding is exactly the “compiler adds nothing over a default” result — the default is just not the cascade.

  3. Drift is where compilation earns its keep — in cost, not accuracy. When quality drifts 0.9, the frozen baselines break the contract (cheapest-adequate 0.778; hand-built cascade 0.797 while costing more, 22.90) or overpay by default (always-strongest 20.00). The compiled plan recompiles to judge and holds the contract at 0.859 — at 20.00, four times the undrifted cost. Cost is the price of holding the contract; that is “worse when they drift” measured in cost and made explicit, never hidden. P3 holds.

  4. explain(plan) states what evidence backed each profile. Every step carries a seed row or ASSUMPTION: linear is seeded by Ch04’s measured tfidf-lr 0.8779 (banking77 test, n=3080); the judge’s curve is an ASSUMPTION because Ch28’s llm_judge cell is NOT_OBSERVED; the cascade’s profile is the Ch30 row verbatim. P4 holds; nothing is invented.

  5. The naive product is never used for selection (P5). The cascade candidate is the supplied end-to-end number, exactly because Ch30 measured the product’s failure.

Wrong / Correct. Wrong: “The compiler chose linear because it is the best model.” Correct: “The compiler chose linear because, under the supplied profiles, it is the cheapest provider whose expected accuracy clears the contract — and the run shows the expert’s cascade misses that same contract. The compiler’s answer is bounded by its profiles; change the profiles and the answer changes (the drift sweep is that change, measured).”

Compile time vs run time, made literal

The plan format keeps the two decision times apart:

  • compile-time decisions — which providers the plan uses, chosen from the supplied profiles (compile_choices); with a single provider there are no runtime rules;
  • runtime decisions — escalation on uncertainty. A cascade plan’s steps carry “answer at threshold tau₀, else escalate”; a single-provider plan carries “always answer”. The schema schemas/plan.schema.json encodes the split.

This is the new bit relative to Selinger: System R’s access-path choice has no runtime phase because SQL queries carry no uncertainty to escalate on. The compile/runtime split is the decision-specific part of the plan.

Prior art, engaged directly

  • Selinger (1979): the shape — cost-based selection from a declarative spec. This chapter is System R for providers; the difference is the runtime escalation half, which has no analogue in access-path selection.
  • Palimpzest (2405.14696): the closest prior art for semantic cost optimization — declarative queries, physical plans over model/hardware/ prompt choices, quality-cost trade-offs. This chapter is narrower: one decision contract, a single plan choice among six families, no workload-wide enumeration, no learned cost models. No novelty is claimed over it.
  • DSPy: optimizes prompts inside a program (compilation as prompt/weight optimization), a different mechanism from plan selection over provider profiles. We compare mechanisms, not performance: neither a DSPy program nor a Palimpzest plan ran here.

(LOTUS and DocETL are not cited in this chapter; their roles here are covered by Palimpzest, which was read. Any citation of them in later chapters stays inside the sections read in Session R.)

What to carry forward

Chapter 34 assembles one program from the surviving constructs. It will need a plan, and the compiler hands it one with the two decision times separated: the program’s provider choices are compile-time facts; its escalation decisions are runtime facts. Chapter 31’s audit lesson applies to the profiles themselves: if a profile was learned from routed traffic, it carries the Chapter 31 bias unless it was audited — the compiler is only as good as the profiles it is given, and the chapter says so.

Close by

What did the compiler learn that a human could not have written down? At target 0.85: that the hand-built cascade misses the contract it was meant to satisfy and that a single adequate classifier beats it on both axes. That number was already in Chapter 30’s rows; the compiler is what read it.

Limitations

  • All numbers are consequences of the supplied profiles: Ch30’s simulation rows seed the curves and the cascade; Ch04’s measured LR row seeds the classifier; the judge and human curves and the retrieval family are ASSUMPTIONS. None of this is evidence about any real router or model.
  • The runtime realization is a SIMULATION over the assumed curves (n=4000); the cascade’s frozen thresholds were fit on the undrifted calibrate stream, then evaluated on drifted curves — a declared construction.
  • Palimpzest and 2403.05821 are read at abstract level; Selinger was verified at DOI level (a classic, not on arXiv). The prior-art comparison is mechanism-level, and no comparative performance is claimed.
  • The cascade-never-chosen result is a property of this profile set; a different cost table (a cheaper judge, a more expensive linear) would put the cascade in the feasible set and the branch would be reachable.
  • The retrieval family is not in the profile set used for the runtime runs; it is covered in Chapter 34 where the program actually retrieves.