Learning the Router
Can a router learn which provider to use, from the outcomes it sees?
Learning the Router
This chapter is a simulation over assumed provider profiles. Every accuracy curve, cost, signal noise and audit fraction is assumed and swept; every number is a consequence of them. No real router (RouteLLM, LinUCB) is measured.
The problem
Chapter 29’s router was explicit rules over metadata. This chapter asks what happens when the router learns the answer from the outcomes it observes. The catch is that what a router observes is decided by the router itself. A router that only sees the outcomes of the providers it selects can be systematically wrong about the providers it avoids — and worse, its mistakes can be self-reinforcing.
What we expect and why
Our starting hypothesis is that a router that only sees the outcomes of providers it selects learns a biased value, and a random audit sample is the mitigation.
Three papers frame it:
-
Mozannar and Sontag (2020) — learning to defer trains a classifier with a reject option; a consistent surrogate for “predict or defer”. The classifier + rejector shape. (Read at abstract level.)
-
Li et al. (2010) — the contextual bandit for personalized recommendation, with a method for offline evaluation on recorded random traffic. The audit sample is exactly that idea: only random traffic gives you an unbiased view of the arms you would not otherwise choose. (Read at abstract level.)
-
Madras et al. (2018) — learning to defer makes a system more accurate and less biased even with inconsistent human decision-makers. Deferral is a real route. (Read at abstract level.)
The build
src/arbiter/learning_router.py simulates two providers, weak and strong, with assumed difficulty-scaled accuracies. Phase 1 routes requests by a noisy difficulty signal and learns each provider’s value only from the items routed to it. Phase 2 deploys the learned value on fresh traffic.
from arbiter.learning_router import LearnedValue
# 1. The bias mechanism: weak only ever sees easy items.
print("1. why the estimate is biased (hand numbers)")
weak = LearnedValue()
for outcome, d in ((True, 0.1), (True, 0.2), (True, 0.3), (False, 0.6)):
weak.correct += outcome
weak.n += 1
print(f" weak saw easy items -> estimate {weak.estimate:.2f} (true marginal 0.68)")
# 2. An audit gives weak a representative sample.
print("2. the audit fixes the estimate")
weak2 = LearnedValue()
for outcome in (True, True, False, False, True, False, False, True):
weak2.correct += outcome
weak2.n += 1
print(f" weak sees mixed items -> estimate {weak2.estimate:.2f}")
1. why the estimate is biased (hand numbers)
weak saw easy items -> estimate 0.75 (true marginal 0.68)
2. the audit fixes the estimate
weak sees mixed items -> estimate 0.50
The sweep
run_ch31.py varies the audit fraction from 0 to 0.25 (n=8,000; curves and competence threshold in every row). results/ch31.jsonl:
| audit | weak estimate | weak marginal | bias | deploy accuracy | regret vs oracle |
|---|---|---|---|---|---|
| 0.00 | 0.864 | 0.673 | +0.190 | 0.680 | 0.238 |
| 0.05 | 0.861 | 0.673 | +0.188 | 0.682 | 0.236 |
| 0.10 | 0.844 | 0.673 | +0.171 | 0.684 | 0.234 |
| 0.25 | 0.791 | 0.673 | +0.117 | 0.949 | −0.031 |
Oracle accuracy is 0.918.
What it says
-
The self-reinforcing failure is real and measurable. With no audit, the router’s learned value of the weak provider is 0.864 against a true marginal of 0.673 — an optimism of +0.19 — because weak only ever sees the easy items the router sends it. Deploying that belief sends almost everything to weak and lands at 0.680 worst-of-sweep accuracy. The router’s own policy created the data that confirmed the policy. P3, the chapter’s point, reproduces.
-
The audit is a priced cure. Every audit point pushes the weak estimate toward its marginal (bias +0.190 → +0.117) and, once enough random traffic has flowed, flips the deployment: at 0.25 the corrected belief drops below the competence threshold and the router sends everything to strong, reaching 0.949. The audit costs accuracy on the audited requests themselves (they are served worse than the routed ones), but it buys an unbiased model. P2 holds; the audit’s benefit is in the model, not in the audited traffic.
-
This deployment is cost-blind. The weak provider costs 1 and the strong 20, but phase 2’s decision is accuracy-only. At audit 0.25 the router collapses to “always strong” — correct but expensive. The cost dimension is deliberately not in the deployment, and no cost saving is claimed; a cost-aware deployment is Chapter 33’s problem.
-
The abrupt threshold flip is an artefact of the single knob. Between audit 0.10 (est 0.844, all weak) and 0.25 (est 0.791, all strong) the deployment jumps. That is the competence threshold 0.20’s own behaviour, and it is reported as such — a swept assumption, not a finding.
Wrong / Correct. Wrong: “A router learns its providers from the outcomes it sees.” Correct: “A router learns every provider it sends traffic to, and learns nothing trustworthy about the ones it avoids. The provider values it learns are a function of its own routing, and the correction — an audit sample — is a tax paid on purpose to see the arms it does not choose.”
What to carry forward
Chapter 30’s cascade needed an uncertainty signal; Chapter 29’s router needed provider profiles. This chapter adds the meta-lesson: if the router learns its own signals, it must audit, because selection confounds its training data. Chapter 33’s compiler will need exactly this: it plans from supplied profiles, and if those profiles are learned from routed traffic, the same audit discipline applies.
Close by
How should a router learn when its own routing determines what it observes? Under this simulation: with an audit sample, priced and swept — and the cost dimension must be in the deployment or the router collapses to always-strong.
Limitations
- SIMULATION over assumed curves, noise and audit fractions;
results/ch31.jsonlismode: simulationwith the assumptions in every row. - The three papers are read at abstract level; full reads were not part of this pass.
- The deployment is accuracy-only; cost-aware routing is deferred to Chapter 33.
- The competence threshold is a single swept knob; the 0.10→0.25 deployment flip is its artefact.
- The audit improves the model; the audited requests themselves are served worse on average — the tax, stated.