The Safe but Useless Model

Explain this chapter with AI

Copy this prompt into ChatGPT, Claude, Gemini, a local model, or another AI.

Apply this chapter with AI

Copy this prompt into ChatGPT, Claude, Gemini, a local model, or another AI.

Chapter 9 changed the unit of evaluation.

Instead of asking whether one answer looks good, we began studying a family of related executions.

That immediately reveals a failure that static evaluation can miss almost completely.

Consider two organizations.

ORGANIZATION A
3 months of runway
falling demand
negative cash flow
credit line nearly exhausted

ORGANIZATION B
5 years of runway
rapidly growing demand
strong margins
large cash reserve

Ask both:

What should management prioritize over the next twelve months?

Now imagine receiving essentially the same answer in both cases:

Invest in innovation.
Improve operational efficiency.
Stay close to customers.
Use data to guide decisions.
Build organizational capabilities.
Balance short-term execution with long-term growth.

There may be no fabricated fact. There may be no contradiction. The advice may be professionally worded and broadly sensible.

And yet something is badly wrong.

The answer barely depends on the problem.

That is the subject of this chapter.

A response can look safe, fluent, grounded, and still be useless because the decisive context never materially entered the decision.

The word safe in the title is deliberately informal. Here it means safe-looking under the static checks developed so far: no obvious fabrication, contradiction, or policy violation. It does not mean that an acceptance policy has certified the recommendation as safe for action.

This failure is broader than hallucination. It is a failure of context utilization.


Where we are

By this point the book can detect several distinct failures:

unsupported semantic extension
structural contradiction
role or polarity failure
instability under irrelevant transformations
insensitivity to declared counterfactual changes

Chapter 10 narrows the last category into an especially important production problem:

What happens when a model repeatedly converges to a polished default even when the task requires different decisions?

The objective is not to reward novelty. It is to test whether decision-relevant context has the leverage the task contract says it should have.


1. The failure is not falsehood

The easiest AI failures to notice are explicit errors:

wrong date
invented citation
reversed relation
unsupported claim
false tool state

The safe-but-useless failure is harder because every individual sentence can survive ordinary inspection.

A static evaluator may report:

Grammatically clear?                    PASS
Obvious factual fabrication?            PASS
Internal contradiction?                 PASS
Obvious policy violation?               PASS
Commonly accepted advice?               PASS

The missing question is relational:

Would the decision have been substantially the same if the important facts had been different?

That question cannot be answered from one completion. We need the response-surface machinery from Chapter 9.


2. Context-insensitive convergence and the default response basin

Call the broader failure:

context-insensitive convergence

A system exhibits context-insensitive convergence when scenarios that the task contract declares decision-distinguishing repeatedly map to the same or nearly the same substantive decision.

The pattern may be:

the same recommendation
the same prioritization
the same ranking
the same trade-off
the same refusal
the same hybrid non-choice

Different inputs do not automatically require different outputs. Two patients may correctly receive the same treatment. Two companies may correctly choose the same liquidity intervention. Two bugs may need the same fix.

So low output diversity is not itself a failure.

Low sensitivity is evidence of genericity only when the oracle says the changed variable should alter the decision.

We will use default response basin as a descriptive behavioral term for a region of a declared scenario family that maps to the same structured decision pattern:

    graph TD
    X1[scenario x1] --> DRB[DEFAULT RESPONSE BASIN]
    X2[scenario x2] --> DRB
    X3[scenario x3] --> DRB
    X4[scenario x4] --> DRB
    X5[scenario x5] --> DRB
  

This is not a claim that the model contains a literal dynamical attractor.

Let:

$$ x_1,x_2,\ldots,x_n $$
be scenarios varying along decision-relevant dimensions, and let:
$$ g(Y) $$
extract the substantive decision from a generated response.

If:

$$ g(Y_{x_1}) \approx \cdots \approx g(Y_{x_n}) $$
while the oracles require different actions, we have observed default-basin behavior.

The text need not be similar.

prioritize customer-centric innovation
invest in differentiated customer value
build innovation capabilities around customer needs

may be three different phrasings of one decision.

Genericity is not textual similarity.


3. Three kinds of collapse

Collapse What remains the same? Is it necessarily a reasoning failure?
Template collapse rhetorical scaffold no
Decision collapse substantive recommendation yes, when oracle requires divergence
Choice avoidance non-committal hybrid yes, when task requires exclusion

A stable template can be useful.

Decision collapse is more serious. Different scenarios receive the same action even though the expected action should change.

Choice avoidance is subtler. The model converts a real trade-off into a balanced-sounding combination:

centralize strategically while decentralizing operationally
pursue radical innovation while maintaining incremental improvement
optimize the short term while investing for the long term

Sometimes a hybrid is genuinely correct. Sometimes combination is a way to evade a required choice.

The task contract must therefore declare the admissible output states:

A
B
HYBRID
ABSTAIN

and whether HYBRID is actually feasible under the stated constraints.

When the task requires exclusion, combination can be evasion.


4. Trendslop is one observed instance

In March 2026, Angelo Romasanta, Llewellyn D. W. Thomas and Natalia Levina reported a closely related pattern in strategic-advice experiments with leading LLMs.[1]

They tested seven recurring business tensions:

exploration vs exploitation
centralization vs decentralization
short-term vs long-term performance
competition vs collaboration
radical vs incremental innovation
differentiation vs commoditization
automation vs augmentation

Across thousands of simulations, the models repeatedly favored fashionable strategic positions rather than context-specific strategic logic. The authors called the pattern trendslop.[1][2]

The important part for this book is the experiment shape:

change scenario context
change framing
repeat across models
        โ†“
observe persistent decision tendencies

That is a response-sensitivity experiment.

The broader technical class is context-insensitive convergence toward a recurring default recommendation basin.

Trendslop is one domain-specific manifestation.


5. Static evaluation can select for genericity

Generic answers can perform well under static evaluation because genericity reduces falsifiable commitment.

Compare:

Cut discretionary R&D immediately and preserve twelve months of payroll.

with:

Balance near-term efficiency with long-term innovation.

The second answer is harder to falsify. It is also harder to act on.

If an evaluator rewards:

fluency
non-toxicity
plausibility
broad helpfulness
absence of explicit factual error

then an evaluation regime can create selection pressure toward answers that minimize contestable commitment.

That does not imply the model consciously optimizes for genericity.

It is a systems-level observation about what kinds of outputs survive the objective.

A response can become safer to score by becoming less useful to decide with.


6. Mechanism hypotheses are not diagnoses

Why might a default basin exist?

Possible mechanisms include:

training-distribution frequency
preference optimization
prompt ambiguity
weak context utilization
evaluator pressure
uncertainty avoidance

These are hypotheses, not conclusions from the behavioral pattern.

It is tempting to blame RLHF or preference tuning directly for the Hybrid Trap. The evidence is not strong enough for that universal claim.

There is, however, relevant evidence that human-preference optimization can create undesirable conditioning behavior. Sharma and colleagues found that human preference data often favored responses matching a user’s stated views, and that optimizing against preference models could sometimes sacrifice truthfulness for sycophancy.[4] Earlier model-written evaluations also found sycophancy and other inverse-scaling behaviors associated with RLHF in some settings.[5]

That supports a narrower conclusion:

Preference objectives can shape which kinds of answers are rewarded, including answers that are agreeable rather than epistemically ideal. It does not establish that preference tuning is the primary cause of context-insensitive convergence.

Sycophancy is useful as a contrast case.

DEFAULT-BASIN FAILURE
underweights decisive local context

SYCOPHANCY
can overweight the user's stated preference
relative to evidence or truth

They are not perfect mirror images, but both show that conditioning can be weighted incorrectly.

Mechanism claims should be tested separately.

For example:

Hypothesis: prompt ambiguity drives hybrid answers
Test: require one mutually exclusive choice

Hypothesis: evaluator pressure drives non-commitment
Test: change judge rubric to reward decision specificity

Hypothesis: generic training frequency dominates
Test: vary framing while preserving the same decisive constraints

Behavior first. Mechanism second.


7. The temperature illusion

When outputs look repetitive, a common response is:

increase temperature
increase top_p

That may increase surface variation, and it may also change the distribution of decisions.

But neither effect proves that the model has become more context-sensitive.

You can easily obtain:

five different phrasings
five different rationales
one repeated recommendation

or, at higher randomness:

more decision variance
without better coupling to the scenario

So the relevant comparison is not:

low-temperature text diversity
vs
high-temperature text diversity

It is:

Does the conditional decision distribution move correctly when the decisive context changes?

Changing decoding parameters can be useful experimentally. It is not a substitute for a context-utilization test.


8. Context ablation should test epistemic behavior, not only decision change

One simple test is context ablation.

Start with:

FULL SCENARIO
company losing money
3 months runway
market contracting
bank covenant near breach

Then remove the decisive details:

ABLATION
A company wants strategic advice.
What should it prioritize?

A weak contract might demand:

decision must change

But that is too rigid. Removing decisive evidence can legitimately produce several outcomes:

same tentative action with lower confidence
more conditional language
request for missing information
abstention
broader recommendation
a different decision

A stronger ablation contract is therefore typed:

context_ablation = {
    "removed_fields": [
        "runway",
        "cash_flow",
        "market_direction",
        "covenant_risk",
    ],
    "expected_relation": {
        "confidence": "decrease",
        "conditionality": "increase",
        "specificity": "not_increase",
        "information_request_or_abstention": "allowed",
        "unqualified_same_decision": "suspicious",
    },
}

This seeds the central distinction of Chapter 11:

GENERIC COLLAPSE
strong evidence exists but is ignored

EPISTEMIC RESTRAINT
evidence is insufficient, so commitment decreases

Those are very different behaviors.


9. Counterfactual inversion measures decision uptake

A stronger test reverses a decisive field:

runway: 18 months โ†’ 3 months
market: growth โ†’ contraction
budget: $10M โ†’ $100k
deadline: 1 year โ†’ 1 week
risk tolerance: high โ†’ near-zero

The perturbation contract specifies what should change:

contract = {
    "name": "runway_inversion",
    "intervention": "18_months -> 3_months",
    "protected_invariants": ["product", "industry", "team_size"],
    "expected_relation": {
        "cash_preservation": "increase",
        "optional_investment": "decrease",
        "planning_horizon": "shorten",
    },
}

Three different observations matter.

Context acknowledgment

Did the response notice the changed fact?

Decision effect

Did the substantive recommendation move?

Directional fidelity

Did it move as the oracle requires?

This separates four cases:

TOTAL INVARIANCE
fact not meaningfully reflected

COSMETIC ADAPTATION
fact mentioned, decision unchanged

ACCIDENTAL / WRONG-DIRECTION RESPONSIVENESS
decision moves, but incorrectly

APPROPRIATE ADAPTATION
decision and rationale move correctly

A sentence such as:

Given the shorter runway, innovation remains essential...

may pass context acknowledgment while failing decision uptake.

That is why keyword use is not evidence of reasoning.


10. Claimed decisive facts are hypotheses, not proof

A useful recommendation should expose its local dependency:

DECISION
DECISIVE FACTS
TRADE-OFF
REJECTED ALTERNATIVE
REVERSAL CONDITION

For example:

recommendation = {
    "decision": "preserve_cash",
    "decisive_facts": [
        "runway_3_months",
        "negative_cash_flow",
        "credit_constraint",
    ],
    "tradeoff": "slower_product_expansion",
    "rejected_alternatives": [
        {
            "option": "large_new_product_bet",
            "reason": "liquidity_constraint",
        }
    ],
    "reversal_condition": "runway_above_18_months_and_positive_cash_flow",
}

But self-reported rationales are not evidence that those facts actually drove the decision.

A model can produce a plausible post-hoc explanation.

So every claimed decisive fact should become an executable hypothesis:

model claims runway=3 months mattered
        โ†“
intervene on runway
        โ†“
regenerate
        โ†“
did the decision move in the declared direction?

This produces a new measurement:

decisive-fact faithfulness

CLAIMED DECISIVE FACT
        โ†“
COUNTERFACTUAL TEST
        โ†“
SUPPORTED ATTRIBUTION
or
UNSUPPORTED RATIONALE

We do not need access to private chain-of-thought. We test the observable policy implied by the explanation.


11. Reversal conditions should be executed

The reversal condition is especially valuable because it is already a falsifiable statement.

Suppose the model says:

I would reverse this recommendation if runway exceeded 18 months
and cash flow became positive.

Construct that scenario.

Regenerate.

If the recommendation does not reverse, then:

declared response policy
โ‰ 
observed response policy

Call this:

reversal-condition fidelity

A good system should not merely state the conditions under which it would change its mind. It should actually change when those conditions are instantiated.

This turns explanation into executable regression testing.


12. Hybrid answers need an explicit exclusion contract

When the task requires a real trade-off, the evaluator needs to know which combinations are impossible or inadmissible.

Let:

$$ \mathcal{O}=\{o_1,\ldots,o_k\} $$
be available options and let an exclusion matrix:
$$ E_{ij}\in\{0,1\} $$
mark pairs that cannot jointly be selected under the task constraints.

For a recommendation set $A(Y)$, a task-specific hybrid conflict rate can be written:

$$ HCR(Y) = \frac{ \sum_{i If there are no relevant exclusion pairs, the metric is `NOT_APPLICABLE`, not zero.

This is not a universal AI score. It is useful only where the domain has an explicit exclusion structure.

A simpler benchmark may just report:

A rate
B rate
HYBRID rate
ABSTAIN rate

under a task contract that declares which states are admissible.

The key distinction remains:

LEGITIMATE HYBRID
both actions are jointly feasible and justified

HYBRID EVASION
the task requires commitment under scarcity,
but the model refuses to choose

13. Decision extraction is another sensor

All of the previous measurements assume that we can map free-form output into a structured decision.

That mapping:

$$ g(Y) $$
is itself fallible.

A decision extractor therefore needs its own contract:

decision_extractor_contract = {
    "input": "free_text_recommendation",
    "output_classes": [
        "PRESERVE_CASH",
        "INVEST_FOR_GROWTH",
        "HYBRID",
        "ABSTAIN",
        "UNKNOWN",
    ],
    "method": "rubric_or_structured_extractor",
    "validation": {
        "human_audit": True,
        "report_accuracy": True,
        "report_agreement": True,
    },
    "known_blind_spots": [
        "implicit_decision",
        "conditional_recommendation",
        "multi-stage_plan",
        "rhetorical_paraphrase",
    ],
}

If $g(Y)$ is wrong, the basin analysis is wrong.

This is another instance of the book’s recurring rule:

Every observable needs a measurement contract.

For production systems, it is often useful to make the generator expose a structured decision payload directly:

from pydantic import BaseModel
from typing import Any

class DecisiveFact(BaseModel):
    variable_name: str
    observed_value: Any
    expected_direction: str | None = None

class ReversalBoundary(BaseModel):
    target_variable: str
    threshold_condition: str
    expected_new_decision: str

class StructuredDecisionPayload(BaseModel):
    primary_decision: str
    decisive_facts: list[DecisiveFact]
    tradeoff: str | None = None
    rejected_alternatives: dict[str, str]
    reversal_boundaries: list[ReversalBoundary]
    confidence: float | None = None
    is_hybrid_choice: bool = False

Structured output does not force correctness. It makes the decision and its claimed dependencies inspectable.


14. Basin occupancy is only a marginal diagnostic

For discrete decisions:

$$ D_1,D_2,\ldots,D_n $$
we can still define dominant model-basin occupancy:
$$ B_{model} = \max_d \frac{1}{n} \sum_i\mathbb{I}(D_i=d), $$
and compare it with:
$$ B_{oracle} = \max_g \frac{1}{n} \sum_i\mathbb{I}(G_i=g). $$
Then:
$$ \Delta B=B_{model}-B_{oracle} $$
measures **excess concentration in the dominant decision**.

But $ฮ”B$ is only a marginal statistic.

Suppose the oracle says:

scenario 1 โ†’ A
scenario 2 โ†’ A
scenario 3 โ†’ B
scenario 4 โ†’ B

and the model says:

scenario 1 โ†’ B
scenario 2 โ†’ B
scenario 3 โ†’ A
scenario 4 โ†’ A

Both distributions have dominant occupancy $0.5$.

The model is still wrong on every case.

So Chapter 10 needs three separate diagnostics:

MARGINAL CONVERGENCE
How concentrated are model decisions?
โ†’ basin occupancy / entropy

SCENARIO AGREEMENT
Did each scenario receive the required decision?
โ†’ oracle agreement

CONTEXT COUPLING
Does changing the decisive context move the decision distribution?
โ†’ paired intervention response

Never substitute the first for the other two.


15. The scenario ร— decision matrix is the primary basin artifact

For scenario families such as:

DISTRESS
STABLE
GROWTH

and decisions:

PRESERVE
OPTIMIZE
INVEST

report the conditional decision matrix:

Scenario Preserve Optimize Invest
Distress 82% 13% 5%
Stable 25% 54% 21%
Growth 11% 19% 70%

A context-insensitive system may instead produce something like:

Scenario Preserve Optimize Invest
Distress 12% 19% 69%
Stable 10% 22% 68%
Growth 11% 20% 69%

These numbers are illustrative, not results from our own experiment.

The real artifact should be compared with the oracle-required matrix and accompanied by uncertainty intervals.

For a deterministic extracted decision, scenario-level agreement is:

$$ A = \frac1n\sum_i\mathbb{I}(D_i=G_i). $$
For repeated stochastic generations, use:
$$ A = \frac1n\sum_iP_{model}(G_i\mid x_i). $$
This is more important than matching the oracle's marginal decision frequencies.

16. Mutual information measures context coupling, not correctness

For controlled scenario classes $X$ and extracted decisions $D$, one can estimate:

$$ I(X;D) = H(D)-H(D\mid X). $$
If the model chooses almost the same decision distribution in every scenario class, $I(X;D)$ will be low.

That makes mutual information a useful context-coupling diagnostic.

One caveat is easy to get wrong. The plug-in estimator of $I(X;D)$ from empirical counts is biased upward, and the bias grows with the number of cells (scenario classes times decision classes) relative to the sample size. With a handful of samples per scenario, a model that ignores context can still post a positive $\hat I(X;D)$ from noise alone. Report a bias-corrected estimate or a permutation baseline โ€” shuffle the scenario labels, recompute, and check that the observed $\hat I$ sits well above that null distribution โ€” before reading any coupling into the number.

One may also compare against the oracle:

$$ \Delta I = I(X;G)-I(X;D). $$
A positive gap can indicate that the model uses less scenario-distinguishing information than the oracle policy requires.

But this is not a correctness metric.

A model can map each scenario class deterministically to the wrong decision and still have high mutual information.

So:

I(X;D)
โ†’ how much decisions depend on context class

oracle agreement
โ†’ whether that dependence is correct

Context utilization is relational. Dependence without directional correctness is not enough.


17. Match the distance to the decision object

Whole-answer cosine similarity is useful for exploration but often wrong for the actual decision.

Use a distance appropriate to the structured unit:

Decision object Example comparison
categorical action exact match / confusion matrix
ranking Kendall’s $\tau$ or rank distance
resource allocation $L_1$ distance over allocation vector
risk level ordinal distance
time horizon interval / ordinal difference
accepted/rejected option exact or set overlap
confidence absolute difference / calibration analysis

The same prose can encode different decisions, and different prose can encode the same decision.

Measure the smallest object that contains the expected effect.


18. Stochastic control uses decision distributions

For a stochastic generator:

$$ Y_x\sim P_\theta(Y\mid x). $$
If the extracted decision is categorical, do not pretend it has a meaningful scalar expectation.

Estimate instead:

$$ P_0(d)=P(D=d\mid x) $$
and:
$$ P_T(d)=P(D=d\mid T(x)). $$
For an expected target decision $d^*$, one simple paired effect is:
$$ \Delta_T(d^*) = P_T(d^*)-P_0(d^*). $$
You may also report distributional movement such as total-variation distance:
$$ TV(P_0,P_T) = \frac12\sum_d|P_0(d)-P_T(d)|. $$
But again, raw movement is not enough.

The experiment must also report whether the mass moved in the oracle-required direction.

For each intervention family report:

samples per condition
within-condition decision entropy
paired decision-change rate
direction-correct rate
scenario-level oracle agreement
confidence intervals
perturbation validity

That separates true context sensitivity from sampling noise.


19. Context length and context utilization are different properties

A common response to generic output is:

Add more context.

Sometimes that works.

But a 5,000-token prompt is not evidence that the model used the five facts that determine the decision.

A direct test compares, for example:

SHORT
3 months runway
negative cash flow

LONG
5,000-token company description
containing the same decisive facts

Then intervene on the same decisive fact in both conditions.

The question is not:

Did the long prompt produce a longer answer?

It is:

Did the decisive variable have more, less, or the same directional effect on the decision?

The issue is not context length.

The issue is decision leverage.


20. Build a domain-specific anti-genericity benchmark

This chapter does not hand you one universal TrendslopScore.

It gives you the specification required to build a falsifiable benchmark for your domain.

A minimum suite might contain:

Decision-relevant inversions

runway: 18 months โ†” 3 months
budget: $10M โ†” $100k
market: growth โ†” contraction
deadline: 1 year โ†” 1 week
risk tolerance: high โ†” near-zero
regulation: permissive โ†” prohibitive

Context ablations

remove budget
remove runway
remove deadline
remove market direction
remove decisive evidence

Invariance controls

paraphrase
format
entity names
irrelevant ordering

Trade-off tests

A
B
HYBRID
ABSTAIN

with an explicit admissibility contract.

Claimed-rationale tests

For every self-reported decisive fact:

intervene
regenerate
test directional response

Reversal tests

For every declared reversal boundary:

instantiate boundary
regenerate
verify reversal

A compact benchmark design could use:

6 inversion types ร— 2 directions = 12 paired conditions
5 ablations
4 invariance controls
3 samples per condition

The exact count depends on cost and domain. What matters is that the transformation family and oracle are frozen before final evaluation.

Required reporting should include:

perturbation validity rate
decision-extractor accuracy/agreement
scenario-level oracle agreement
context acknowledgment rate
paired decision-change rate
direction-correct rate
within-condition entropy
basin occupancy
excess basin occupancy
hybrid/non-choice rate
decisive-fact faithfulness
reversal-condition fidelity

This is enough to turn:

the answer feels generic

into a reproducible measurement problem.

Our earlier internal exploration of the Trendslop hypothesis supplied an initial perturbation taxonomy, semantic-divergence sketches, and response-manifold intuition. It did not produce the controlled book-owned benchmark required by the stricter Chapters 9โ€“10 protocol.

That is not a reason to invent a number.

It is a specification for the next experiment.


21. Production reality: use the surface selectively

Full counterfactual evaluation is expensive.

A base prompt plus five perturbations and three samples per condition already requires eighteen generations.

That may be appropriate for a benchmark. It may be absurd as a synchronous gate.

Separate:

CHEAP SYNCHRONOUS
containment
structured constraints
runtime verification

SHADOW
counterfactual suites on sampled traffic

TEMPLATE / WORKLOAD LEVEL
representative prompt-family testing

RISK TRIGGERED
high-impact decisions
low decision specificity
high uncertainty
novel scenario family

OFFLINE REGRESSION
known counterfactual families after changes

MODEL / CONFIG RELEASE GATE
full sensitivity suite before promoting
model, prompt, retrieval, or policy versions

A compact asynchronous shadow check can be as simple as:

import asyncio

async def shadow_counterfactual_check(
    generate,
    extract_decision,
    base_prompt,
    counterfactual_prompt,
    expected_relation,
):
    base_text, cf_text = await asyncio.gather(
        generate(base_prompt),
        generate(counterfactual_prompt),
    )

    base_decision = extract_decision(base_text)
    cf_decision = extract_decision(cf_text)

    return {
        "base_decision": base_decision,
        "counterfactual_decision": cf_decision,
        "relation_passed": expected_relation(
            base_decision,
            cf_decision,
        ),
    }

This is an orchestration sketch, not a production-complete evaluator. Real systems still need repeated sampling, perturbation validation, extractor validation, tracing, and uncertainty estimates.


22. Repair the system, not only the prose

Once context-insensitive convergence is detected, several interventions are possible.

Force explicit trade-offs

Require:

one primary recommendation
one rejected alternative
reason for rejection

Require structured decision output

A typed payload makes:

decision
decisive facts
tradeoff
rejected alternatives
reversal boundaries

observable.

This can act as a useful mechanical wedge against vague hybrid output, but it does not guarantee correct reasoning.

Generate alternatives before selecting

Use the model to widen the option set, then evaluate alternatives under explicit constraints.

Counterfactually verify the recommendation

Intervene on a claimed decisive fact and test whether the decision moves correctly.

Separate proposal from policy

Let the model propose possibilities. Let a later policy or optimization layer decide what is admissible.

Escalate ambiguous trade-offs

If the available evidence cannot distinguish the options, the right response may be review or abstention rather than a forced answer.

The last intervention leads directly to Chapter 11.


23. The safe-but-useless signature in the reliability record

The diagnostic record can now expose a failure that static hallucination checks miss:

containment                 = PASS
relation_fidelity           = PASS
polarity_fidelity           = PASS
provenance                  = VERIFIED
invariance_controls         = PASS
context_acknowledgment      = PASS
counterfactual_decision     = FAIL
directional_fidelity        = FAIL
context_ablation            = FAIL
decisive_fact_faithfulness  = FAIL
reversal_condition_fidelity = FAIL
decision_specificity        = LOW

Nothing here says:

hallucination detected

Yet the output should not be trusted as a context-specific recommendation.

The pattern is:

containment = PASS
constraint fidelity = PASS
context sensitivity = FAIL

This is the signature Chapter 9 prepared us to observe.


24. What you should now be able to answer

After this chapter, you should be able to explain:

  1. Why different wording is not evidence of different decisions.
  2. Why basin occupancy is useful but cannot replace scenario-level oracle agreement.
  3. How context ablation differs from counterfactual inversion.
  4. Why mentioning a decisive fact does not prove the fact influenced the recommendation.
  5. How to test a model’s stated reversal condition.
  6. When a hybrid answer is a legitimate composition and when it is choice avoidance.
  7. Why higher temperature does not establish greater context sensitivity.
  8. Why more context tokens do not prove more context utilization.

Exercises

Exercise 1 โ€” Build a two-basin test

Construct ten distress scenarios and ten growth scenarios with clear oracle decisions. Generate multiple responses per scenario, extract one structured decision, and report:

scenario ร— decision matrix
oracle agreement
basin occupancy
hybrid rate
direction-correct counterfactual rate

Exercise 2 โ€” Test a claimed decisive fact

Ask a model to provide a recommendation plus the three facts it considers decisive. Change one claimed fact while holding the others fixed. Does the recommendation move in the predicted direction?

Exercise 3 โ€” Execute a reversal condition

Require the model to state what condition would reverse its recommendation. Construct that condition and rerun the model. Record reversal-condition fidelity.

Exercise 4 โ€” Separate diversity from context use

Run the same scenario at several decoding temperatures. Then run two decision-distinguishing scenarios at one fixed temperature. Compare lexical diversity with decision-distribution movement. Which change actually tracks context?


25. The deeper lesson

Generic answers are attractive because they survive many evaluators. They are hard to falsify, sound reasonable, avoid controversial choices, and often contain familiar best practices.

But decision support is not a contest in producing statements that are difficult to disagree with.

It is a task of conditioning recommendations on constraints.

The relevant distinction is:

GOOD GENERAL ADVICE
can be broadly useful across many cases

GENERIC COLLAPSE
persists even when the task contract
says the decision should change

So the engineering rule is:

Do not ask only whether an answer is reasonable. Ask which facts made this answer different from the one the system would have produced anyway.

And then test those facts.

A fact is not demonstrated to be decision-relevant because the model mentions it.

A fact is demonstrated behaviorally when changing it changes the decision in the relation the task requires.

If we cannot observe that dependency, we may not have observed context-sensitive reasoning.

We may have observed a polished default.


Research roots

  1. Angelo Romasanta, Llewellyn D. W. Thomas and Natalia Levina, “Researchers Asked LLMs for Strategic Advice. They Got ‘Trendslop’ in Return,” Harvard Business Review, March 16, 2026. Reports persistent strategic biases across seven core business tensions and warns that LLM recommendations can follow fashionable managerial patterns rather than context-specific strategic logic. https://hbr.org/2026/03/researchers-asked-llms-for-strategic-advice-they-got-trendslop-in-return

  2. Natalia Levina, “Researchers Asked LLMs for Strategic Advice. They Got ‘Trendslop’ in Return,” NYU Stern Research Highlight, March 16, 2026. Summarizes the study as thousands of simulations in which leading LLMs repeatedly selected trendy strategic options across varied contexts. https://www.stern.nyu.edu/experience-stern/faculty-research/research-highlights/researchers-asked-llms-strategic-advice-they-got-trendslop-return

  3. Harvard Business Review, product summary for Romasanta, Thomas and Levina, “Researchers Asked LLMs for Strategic Advice. They Got ‘Trendslop’ in Return,” 2026. Practical guidance includes using LLMs to expand options rather than make final strategic choices, counteracting biases, avoiding the hybrid trap, and not assuming that context alone removes the bias. https://store.hbr.org/product/researchers-asked-llms-for-strategic-advice-they-got-trendslop-in-return/H093GG

  4. Mrinank Sharma et al., “Towards Understanding Sycophancy in Language Models,” 2023. Finds sycophancy across several state-of-the-art assistants, shows that human preference data can favor answers matching user views, and reports that preference-model optimization can sometimes trade truthfulness for sycophancy. https://arxiv.org/abs/2310.13548

  5. Ethan Perez et al., “Discovering Language Model Behaviors with Model-Written Evaluations,” Findings of ACL 2023, pp. 13387โ€“13434. Uses model-written evaluations to identify behaviors including sycophancy and reports examples where RLHF exacerbated undesirable behaviors. https://aclanthology.org/2023.findings-acl.847/

Next: Knowing When Not to Answer

The safe-but-useless model exposes one more ambiguity.

Suppose the model does not adapt strongly to the context.

Why?

Possibility one:

the model ignored decisive facts

But there is another possibility:

the evidence genuinely does not justify a decisive answer

In that case, refusing to make a sharp recommendation may be correct.

We therefore need another question:

Does the system have enough evidence to answer at all?

That is not containment. It is not consistency. It is not sensitivity.

It is epistemic adequacy.

Chapter 11 turns abstention from an embarrassing non-answer into a measurable system capability.