← Jev From First Principles

Decision Types

Which output types are genuinely different decision semantics, and which are only representations of Choice plus post-processing?

Decision Types

The problem

You have a decision to make. The answer could be one of several options, or a number on a scale, or a preference between two things, or a set of options. The vendor’s contract gives you three question types: Choice, Score, and Noul. But the programming language you are building needs to know which of these are genuinely different semantics and which are just representations of the same underlying thing.

This chapter asks: how many decision types does the language need?

The answer matters because every type you add to the language is a commitment. It needs syntax, validation, calibration semantics, abstention rules, and a story for what it means. If a type is just a representation of another type, it does not deserve its own syntax — it deserves a library function.

What we expect and why

Our starting hypothesis is that Choice and Preference are primitive, and everything else is derivable:

  • Binary is a 2-option Choice.
  • Noul is a Binary with a calibrated probability.
  • Score is a Choice over ordinal levels, but the ordinal structure is lost when treated as nominal categories.
  • Probability is a Noul with calibration.
  • Set and Optional are type constructors over Choice, not independent semantics.
  • Preference is different because pairwise comparison is a different elicitation mode with different noise properties.

Three papers inform this expectation:

  1. Cao et al. (2019) — CORAL shows that ordinal regression needs rank-consistency guarantees that nominal classification lacks. If you treat ordinal levels as nominal categories, you lose the ordering information. This is the argument against Score collapsing to Choice without loss.

  2. Christiano et al. (2017) — Learning from pairwise human preferences shows that comparative decisions (A vs B) are a natural primitive. The Bradley-Terry model converts comparisons to rewards, but the elicitation mode is different from scoring: humans are more consistent at comparing than at assigning absolute numbers.

  3. Tian et al. (2023) — Just Ask for Calibration shows that numeric scores are often uncalibrated regardless of the output type. Verbalized confidences frequently outperform log probabilities. This challenges the idea that a Score of 0.63 means anything stable across providers.

The build

We implement each type as a collapse function to Choice plus post-processing. The collapse is information-preserving if you can round-trip without loss; it is lossy if the original type’s invariants do not survive.

The code is in src/arbiter/types.py. It is model-free: no model, no inference, no downloaded data. The tests are in tests/types/test_types.py (39 tests, all passing).

Binary collapses to Choice

Binary is the simplest case. A yes/no decision is exactly a 2-option Choice:

from arbiter.types import (
    analyze_binary,
    analyze_noul,
    analyze_optional,
    analyze_preference,
    analyze_score,
    analyze_set,
    binary_to_choice,
    choice_to_binary,
    minimal_type_set,
    noul_to_choice,
    noul_to_choice_with_prob,
    preference_to_choice,
    preference_to_choice_with_margin,
    score_agreement,
    score_calibration_error,
    score_to_choice,
    score_to_choice_with_ordinal,
)
    # 1. Binary collapses to Choice with zero information loss
    print("1. Binary collapses to Choice (zero information loss)")
    for val in (True, False):
        choice = binary_to_choice(val)
        back = choice_to_binary(choice)
        print(f"   {val} -> {choice!r} -> {back} (round-trip: {'OK' if back == val else 'FAIL'})")

The round-trip is exact. Binary collapses to Choice with zero information loss.

Noul collapses to Choice

Noul (yes/no probability) collapses to Binary, which collapses to Choice. The probability is kept alongside the choice:

    # 2. Noul collapses to Choice with zero information loss
    print("2. Noul collapses to Choice (zero information loss)")
    for noul in (0.9, 0.5, 0.1):
        choice, prob = noul_to_choice_with_prob(noul)
        print(f"   noul={noul} -> choice={choice!r}, prob={prob} (kept)")

The probability is the calibrated confidence. Noul collapses to Choice with zero information loss if the probability is kept.

Score collapses to Choice with information loss

Score is where the collapse becomes lossy. A 5-level score treated as nominal categories loses the ordering and distance between levels:

    # 3. Score collapses to Choice with information loss
    print("3. Score collapses to Choice (information loss: ordering and distance)")
    for score in (0.0, 0.25, 0.5, 0.75, 1.0):
        level = score_to_choice(score, levels=5)
        print(f"   score={score} -> level {level} (ordering lost)")
    level, probs = score_to_choice_with_ordinal(0.5, levels=5)
    print(f"   score=0.5 with ordinal probs: level {level}, probs={[round(p, 3) for p in probs]}")

The second form keeps a probability distribution over levels, which preserves the ordering in the representation. But the Choice semantics do not enforce the ordering — a provider could return any distribution, and the collapse would not know the difference.

This is the CORAL argument: ordinal outputs need consistency guarantees that nominal classification lacks. The K-1 binary classifiers in a naive ordinal regression can disagree; CORAL fixes this with shared weights and ordered biases. Our collapse has the same problem: treating levels as nominal categories loses the rank-consistency guarantee.

Preference collapses to Choice

Preference (A vs B) collapses to a 2-option Choice. The winner is preserved; the margin is lost unless separately recorded:

    # 4. Preference collapses to Choice with zero information loss in the winner
    print("4. Preference collapses to Choice (zero information loss in winner)")
    for winner in ("A", "B"):
        choice = preference_to_choice(winner, "A", "B")
        print(f"   {winner} vs {'B' if winner == 'A' else 'A'} -> {choice!r}")
    choice, margin = preference_to_choice_with_margin("A", "A", "B", 0.8)
    print(f"   A preferred with margin 0.8 -> choice={choice!r}, margin={margin} (kept)")

The margin is the calibrated preference probability. Preference collapses to Choice with zero information loss in the winner, but the margin must be kept separately.

The reason Preference is primitive is not the output shape — it is the elicitation mode. Christiano et al. showed that humans are more consistent at comparing than at assigning absolute numbers. The Bradley-Terry model converts comparisons to rewards, but the comparison itself is the primitive.

Set and Optional are type constructors

Set and Optional are not decision types. They are type constructors over Choice:

    # 5. Set and Optional are type constructors, not independent semantics
    print("5. Set and Optional are type constructors over Choice")
    for name, analyzer in (("Set", analyze_set), ("Optional", analyze_optional)):
        result = analyzer()
        print(f"   {name}: collapses_without_loss={result.collapses_without_loss}")

They do not need their own syntax. They need a type constructor in the type system.

Cross-provider agreement on scores

    # 6. Cross-provider agreement on scores
    print("6. Cross-provider agreement on scores")
    for a, b in ((0.63, 0.65), (0.63, 0.80)):
        agree = score_agreement(a, b, tolerance=0.1)
        print(f"   provider A={a}, provider B={b} -> agree={agree}")

Score calibration error

    # 7. Score calibration error
    print("7. Score calibration error (ILLUSTRATIVE numbers)")
    good_scores = [0.1, 0.3, 0.7, 0.9]
    good_outcomes = [False, False, True, True]
    good_ece = score_calibration_error(good_scores, good_outcomes)
    print(f"   well-calibrated: ECE={good_ece:.3f}")
    bad_scores = [0.9, 0.8, 0.7, 0.6]
    bad_outcomes = [False, False, True, True]
    bad_ece = score_calibration_error(bad_scores, bad_outcomes)
    print(f"   poorly calibrated: ECE={bad_ece:.3f}")

Minimal type set

    # 8. Minimal type set
    print("8. Minimal type set")
    for name, desc in minimal_type_set().items():
        print(f"   {name}: {desc}")

The result

The walkthrough (examples/ch19-decision-types/walkthrough_ch19.py) runs all collapse functions on hand-checkable inputs. The output shows:

1. Binary collapses to Choice (zero information loss)
   True -> 'yes' -> True (round-trip: OK)
   False -> 'no' -> False (round-trip: OK)
2. Noul collapses to Choice (zero information loss)
   noul=0.9 -> choice='yes', prob=0.9 (kept)
   noul=0.5 -> choice='yes', prob=0.5 (kept)
   noul=0.1 -> choice='no', prob=0.1 (kept)
3. Score collapses to Choice (information loss: ordering and distance)
   score=0.0 -> level 0 (ordering lost)
   score=0.25 -> level 1 (ordering lost)
   score=0.5 -> level 2 (ordering lost)
   score=0.75 -> level 3 (ordering lost)
   score=1.0 -> level 4 (ordering lost)
   score=0.5 with ordinal probs: level 2, probs=[0.067, 0.183, 0.498, 0.183, 0.067]
4. Preference collapses to Choice (zero information loss in winner)
   A vs B -> 'A'
   B vs A -> 'B'
   A preferred with margin 0.8 -> choice='A', margin=0.8 (kept)
5. Set and Optional are type constructors over Choice
   Set: collapses_without_loss=True
   Optional: collapses_without_loss=True
6. Cross-provider agreement on scores
   provider A=0.63, provider B=0.65 -> agree=True
   provider A=0.63, provider B=0.8 -> agree=False
7. Score calibration error (ILLUSTRATIVE numbers)
   well-calibrated: ECE=0.200
   poorly calibrated: ECE=0.600
8. Minimal type set
   Choice: primitive: the base decision type
   Preference: primitive: pairwise comparison is a different elicitation mode
   Binary: constructor: 2-option Choice
   Noul: constructor: Binary with calibration
   Score: constructor: Choice with ordinal post-processing (information loss)
   Probability: constructor: Noul with calibration
   Set: constructor: Choice over power set
   Optional: constructor: Choice over options plus None

The minimal type set is:

Type Status Reason
Choice primitive the base decision type
Preference primitive pairwise comparison is a different elicitation mode
Binary constructor 2-option Choice
Noul constructor Binary with calibration
Score constructor Choice with ordinal post-processing (information loss)
Probability constructor Noul with calibration
Set constructor Choice over power set
Optional constructor Choice over options plus None

What surprised us

  1. Score is the only type that loses information under collapse. Binary, Noul, and Preference all round-trip exactly. Score does not: the ordering and distance between levels are lost when treated as nominal categories. This is the CORAL argument applied to type systems.

  2. The vendor’s Noul is a yes/no probability, nothing more. It is not a separate semantic type. It is a Binary with a calibrated probability. The vendor’s contract exposes it as a separate question type, but that is a packaging decision, not a semantic one.

  3. Cross-provider agreement on scores is not guaranteed. Two providers can give scores of 0.63 and 0.65 for the same input, which looks like agreement. But without calibration, the scores may not mean the same thing. Tian et al. showed that numeric scores are often uncalibrated regardless of the output type.

  4. Preference is primitive because of the elicitation mode, not the output shape. The output is just a 2-option Choice. But the process of getting there — pairwise comparison — has different noise properties than scoring. Christiano et al. showed that humans are more consistent at comparing than at assigning absolute numbers.

The distinction this chapter keeps

Primitive vs constructor. A primitive type has semantics that cannot be reduced to another type without loss. A constructor is a representation that can be reduced to a primitive with post-processing. The language should expose primitives as syntax and constructors as library functions.

Elicitation mode vs output shape. Two types can have the same output shape but different elicitation modes. Preference and Binary both produce a 2-option Choice, but the process of getting there is different. The language should expose the elicitation mode as primitive if it has different noise properties.

What to carry forward

The minimal type set is Choice and Preference as primitives, with Binary, Noul, Score, Probability, Set, and Optional as constructors. This means:

  • The language needs syntax for Choice and Preference.
  • The language needs library functions for Binary, Noul, Score, Probability, Set, and Optional.
  • Score needs ordinal post-processing that preserves ordering, not just a mapping to nominal categories.
  • Calibration is a property of the provider, not the type. A Score of 0.63 is not inherently calibrated.

Close by

Which Jev primitives survived? The vendor’s contract exposes Choice, Score, and Noul as separate question types. Our analysis suggests that only Choice is primitive. Score is a constructor with information loss. Noul is a constructor without information loss. The vendor’s packaging is not the same as the semantics.

The next chapter asks: given that the decision layer must be fed, who decides what it reads?

Limitations

  • This is a model-free type-system analysis. No model was run. The collapse functions are pure Python on hand-checkable inputs.
  • The calibration error numbers in the walkthrough are ILLUSTRATIVE, not measured from a real provider.
  • The cross-provider agreement check is a simple tolerance check, not a statistical test.
  • The three required papers are read in full, but the analysis is our interpretation, not a replication of their experiments.