Decision Types
Which output types are genuinely different decision semantics, and which are only representations of Choice plus post-processing?
Decision Types
The problem
You have a decision to make. The answer could be one of several options, or a number on a scale, or a preference between two things, or a set of options. The vendor’s contract gives you three question types: Choice, Score, and Noul. But the programming language you are building needs to know which of these are genuinely different semantics and which are just representations of the same underlying thing.
This chapter asks: how many decision types does the language need?
The answer matters because every type you add to the language is a commitment. It needs syntax, validation, calibration semantics, abstention rules, and a story for what it means. If a type is just a representation of another type, it does not deserve its own syntax — it deserves a library function.
What we expect and why
Our starting hypothesis is that Choice and Preference are primitive, and everything else is derivable:
- Binary is a 2-option Choice.
- Noul is a Binary with a calibrated probability.
- Score is a Choice over ordinal levels, but the ordinal structure is lost when treated as nominal categories.
- Probability is a Noul with calibration.
- Set and Optional are type constructors over Choice, not independent semantics.
- Preference is different because pairwise comparison is a different elicitation mode with different noise properties.
Three papers inform this expectation:
-
Cao et al. (2019) — CORAL shows that ordinal regression needs rank-consistency guarantees that nominal classification lacks. If you treat ordinal levels as nominal categories, you lose the ordering information. This is the argument against Score collapsing to Choice without loss.
-
Christiano et al. (2017) — Learning from pairwise human preferences shows that comparative decisions (A vs B) are a natural primitive. The Bradley-Terry model converts comparisons to rewards, but the elicitation mode is different from scoring: humans are more consistent at comparing than at assigning absolute numbers.
-
Tian et al. (2023) — Just Ask for Calibration shows that numeric scores are often uncalibrated regardless of the output type. Verbalized confidences frequently outperform log probabilities. This challenges the idea that a Score of 0.63 means anything stable across providers.
The build
We implement each type as a collapse function to Choice plus post-processing. The collapse is information-preserving if you can round-trip without loss; it is lossy if the original type’s invariants do not survive.
The code is in src/arbiter/types.py. It is model-free: no model, no inference, no downloaded data. The tests are in tests/types/test_types.py (39 tests, all passing).
Binary collapses to Choice
Binary is the simplest case. A yes/no decision is exactly a 2-option Choice:
from arbiter.types import (
analyze_binary,
analyze_noul,
analyze_optional,
analyze_preference,
analyze_score,
analyze_set,
binary_to_choice,
choice_to_binary,
minimal_type_set,
noul_to_choice,
noul_to_choice_with_prob,
preference_to_choice,
preference_to_choice_with_margin,
score_agreement,
score_calibration_error,
score_to_choice,
score_to_choice_with_ordinal,
)
# 1. Binary collapses to Choice with zero information loss
print("1. Binary collapses to Choice (zero information loss)")
for val in (True, False):
choice = binary_to_choice(val)
back = choice_to_binary(choice)
print(f" {val} -> {choice!r} -> {back} (round-trip: {'OK' if back == val else 'FAIL'})")
The round-trip is exact. Binary collapses to Choice with zero information loss.
Noul collapses to Choice
Noul (yes/no probability) collapses to Binary, which collapses to Choice. The probability is kept alongside the choice:
# 2. Noul collapses to Choice with zero information loss
print("2. Noul collapses to Choice (zero information loss)")
for noul in (0.9, 0.5, 0.1):
choice, prob = noul_to_choice_with_prob(noul)
print(f" noul={noul} -> choice={choice!r}, prob={prob} (kept)")
The probability is the calibrated confidence. Noul collapses to Choice with zero information loss if the probability is kept.
Score collapses to Choice with information loss
Score is where the collapse becomes lossy. A 5-level score treated as nominal categories loses the ordering and distance between levels:
# 3. Score collapses to Choice with information loss
print("3. Score collapses to Choice (information loss: ordering and distance)")
for score in (0.0, 0.25, 0.5, 0.75, 1.0):
level = score_to_choice(score, levels=5)
print(f" score={score} -> level {level} (ordering lost)")
level, probs = score_to_choice_with_ordinal(0.5, levels=5)
print(f" score=0.5 with ordinal probs: level {level}, probs={[round(p, 3) for p in probs]}")
The second form keeps a probability distribution over levels, which preserves the ordering in the representation. But the Choice semantics do not enforce the ordering — a provider could return any distribution, and the collapse would not know the difference.
This is the CORAL argument: ordinal outputs need consistency guarantees that nominal classification lacks. The K-1 binary classifiers in a naive ordinal regression can disagree; CORAL fixes this with shared weights and ordered biases. Our collapse has the same problem: treating levels as nominal categories loses the rank-consistency guarantee.
Preference collapses to Choice
Preference (A vs B) collapses to a 2-option Choice. The winner is preserved; the margin is lost unless separately recorded:
# 4. Preference collapses to Choice with zero information loss in the winner
print("4. Preference collapses to Choice (zero information loss in winner)")
for winner in ("A", "B"):
choice = preference_to_choice(winner, "A", "B")
print(f" {winner} vs {'B' if winner == 'A' else 'A'} -> {choice!r}")
choice, margin = preference_to_choice_with_margin("A", "A", "B", 0.8)
print(f" A preferred with margin 0.8 -> choice={choice!r}, margin={margin} (kept)")
The margin is the calibrated preference probability. Preference collapses to Choice with zero information loss in the winner, but the margin must be kept separately.
The reason Preference is primitive is not the output shape — it is the elicitation mode. Christiano et al. showed that humans are more consistent at comparing than at assigning absolute numbers. The Bradley-Terry model converts comparisons to rewards, but the comparison itself is the primitive.
Set and Optional are type constructors
Set and Optional are not decision types. They are type constructors over Choice:
# 5. Set and Optional are type constructors, not independent semantics
print("5. Set and Optional are type constructors over Choice")
for name, analyzer in (("Set", analyze_set), ("Optional", analyze_optional)):
result = analyzer()
print(f" {name}: collapses_without_loss={result.collapses_without_loss}")
They do not need their own syntax. They need a type constructor in the type system.
Cross-provider agreement on scores
# 6. Cross-provider agreement on scores
print("6. Cross-provider agreement on scores")
for a, b in ((0.63, 0.65), (0.63, 0.80)):
agree = score_agreement(a, b, tolerance=0.1)
print(f" provider A={a}, provider B={b} -> agree={agree}")
Score calibration error
# 7. Score calibration error
print("7. Score calibration error (ILLUSTRATIVE numbers)")
good_scores = [0.1, 0.3, 0.7, 0.9]
good_outcomes = [False, False, True, True]
good_ece = score_calibration_error(good_scores, good_outcomes)
print(f" well-calibrated: ECE={good_ece:.3f}")
bad_scores = [0.9, 0.8, 0.7, 0.6]
bad_outcomes = [False, False, True, True]
bad_ece = score_calibration_error(bad_scores, bad_outcomes)
print(f" poorly calibrated: ECE={bad_ece:.3f}")
Minimal type set
# 8. Minimal type set
print("8. Minimal type set")
for name, desc in minimal_type_set().items():
print(f" {name}: {desc}")
The result
The walkthrough (examples/ch19-decision-types/walkthrough_ch19.py) runs all collapse functions on hand-checkable inputs. The output shows:
1. Binary collapses to Choice (zero information loss)
True -> 'yes' -> True (round-trip: OK)
False -> 'no' -> False (round-trip: OK)
2. Noul collapses to Choice (zero information loss)
noul=0.9 -> choice='yes', prob=0.9 (kept)
noul=0.5 -> choice='yes', prob=0.5 (kept)
noul=0.1 -> choice='no', prob=0.1 (kept)
3. Score collapses to Choice (information loss: ordering and distance)
score=0.0 -> level 0 (ordering lost)
score=0.25 -> level 1 (ordering lost)
score=0.5 -> level 2 (ordering lost)
score=0.75 -> level 3 (ordering lost)
score=1.0 -> level 4 (ordering lost)
score=0.5 with ordinal probs: level 2, probs=[0.067, 0.183, 0.498, 0.183, 0.067]
4. Preference collapses to Choice (zero information loss in winner)
A vs B -> 'A'
B vs A -> 'B'
A preferred with margin 0.8 -> choice='A', margin=0.8 (kept)
5. Set and Optional are type constructors over Choice
Set: collapses_without_loss=True
Optional: collapses_without_loss=True
6. Cross-provider agreement on scores
provider A=0.63, provider B=0.65 -> agree=True
provider A=0.63, provider B=0.8 -> agree=False
7. Score calibration error (ILLUSTRATIVE numbers)
well-calibrated: ECE=0.200
poorly calibrated: ECE=0.600
8. Minimal type set
Choice: primitive: the base decision type
Preference: primitive: pairwise comparison is a different elicitation mode
Binary: constructor: 2-option Choice
Noul: constructor: Binary with calibration
Score: constructor: Choice with ordinal post-processing (information loss)
Probability: constructor: Noul with calibration
Set: constructor: Choice over power set
Optional: constructor: Choice over options plus None
The minimal type set is:
| Type | Status | Reason |
|---|---|---|
| Choice | primitive | the base decision type |
| Preference | primitive | pairwise comparison is a different elicitation mode |
| Binary | constructor | 2-option Choice |
| Noul | constructor | Binary with calibration |
| Score | constructor | Choice with ordinal post-processing (information loss) |
| Probability | constructor | Noul with calibration |
| Set | constructor | Choice over power set |
| Optional | constructor | Choice over options plus None |
What surprised us
-
Score is the only type that loses information under collapse. Binary, Noul, and Preference all round-trip exactly. Score does not: the ordering and distance between levels are lost when treated as nominal categories. This is the CORAL argument applied to type systems.
-
The vendor’s Noul is a yes/no probability, nothing more. It is not a separate semantic type. It is a Binary with a calibrated probability. The vendor’s contract exposes it as a separate question type, but that is a packaging decision, not a semantic one.
-
Cross-provider agreement on scores is not guaranteed. Two providers can give scores of 0.63 and 0.65 for the same input, which looks like agreement. But without calibration, the scores may not mean the same thing. Tian et al. showed that numeric scores are often uncalibrated regardless of the output type.
-
Preference is primitive because of the elicitation mode, not the output shape. The output is just a 2-option Choice. But the process of getting there — pairwise comparison — has different noise properties than scoring. Christiano et al. showed that humans are more consistent at comparing than at assigning absolute numbers.
The distinction this chapter keeps
Primitive vs constructor. A primitive type has semantics that cannot be reduced to another type without loss. A constructor is a representation that can be reduced to a primitive with post-processing. The language should expose primitives as syntax and constructors as library functions.
Elicitation mode vs output shape. Two types can have the same output shape but different elicitation modes. Preference and Binary both produce a 2-option Choice, but the process of getting there is different. The language should expose the elicitation mode as primitive if it has different noise properties.
What to carry forward
The minimal type set is Choice and Preference as primitives, with Binary, Noul, Score, Probability, Set, and Optional as constructors. This means:
- The language needs syntax for Choice and Preference.
- The language needs library functions for Binary, Noul, Score, Probability, Set, and Optional.
- Score needs ordinal post-processing that preserves ordering, not just a mapping to nominal categories.
- Calibration is a property of the provider, not the type. A Score of 0.63 is not inherently calibrated.
Close by
Which Jev primitives survived? The vendor’s contract exposes Choice, Score, and Noul as separate question types. Our analysis suggests that only Choice is primitive. Score is a constructor with information loss. Noul is a constructor without information loss. The vendor’s packaging is not the same as the semantics.
The next chapter asks: given that the decision layer must be fed, who decides what it reads?
Limitations
- This is a model-free type-system analysis. No model was run. The collapse functions are pure Python on hand-checkable inputs.
- The calibration error numbers in the walkthrough are ILLUSTRATIVE, not measured from a real provider.
- The cross-provider agreement check is a simple tolerance check, not a statistical test.
- The three required papers are read in full, but the analysis is our interpretation, not a replication of their experiments.