← Language From First Principles

Too Much Information

Reframe information abundance as an attention-allocation problem.

The Semantic Browser works. On any page you open, it chooses a representation, enriches it, types its relationships, and gates everything through preservation evidence. Now count what it never touches: the forty hours of conference video in your subscriptions, the papers published this week in your field, the podcasts queued behind the podcasts, the threads, the advisories, the second-order citations of everything you read. The browser perfects the encounter with information already in front of you. It does nothing for the information you never reach — which is nearly all of it.

That is Part II’s stricter problem, and Chapter 11 states it as constraint rather than complaint:

Even a perfect Semantic Browser cannot help with information the user never reaches.

The chapter’s job is to turn that sentence into machinery: a resource model in which attention is a budget, supply exceeds it structurally, and selection becomes unavoidable. It does not solve which items matter — importance ranking belongs to Chapters 13 through 16, novelty to 15, timing to 19. Here: the constraint, the funnel, the unit of value, and the curve that pressures everything downstream.

The constraint: supply structurally exceeds capacity

information supply  >>  attention capacity

This is not a mood but an arithmetic fact about the reader’s world. Supply grows with every subscribed channel, every publishing researcher, every recorded talk — multiplicatively, continuously, without consulting the reader’s calendar. Capacity is bounded by hours, seriality, and fatigue: one person, one stream of attention at a time (Simon’s 1971 point, content-read for this book: a wealth of information creates a poverty of attention, and organisations — now individuals — must allocate it). The gap between the two is not closable by reading faster, and Part I’s machinery, for all its power, operates exclusively downstream of the gap: it improves the page you opened, not the forty you didn’t.

The interesting quantity is therefore not compression ratio — how many hours became how many minutes — but something closer to:

value gained per unit of attention spent

Stated deliberately without collapsing it into one universal score. Value differs by task (the implementer and the browser want different things from the same hour), attention differs by state (five focused minutes are not five distracted ones), and any single number would smuggle in exactly the objective-importance claim this chapter refuses. The ratio is a framing device: every Part II mechanism will be scored by what it puts into the numerator per unit of the denominator, measured separately per task, with the components reported rather than averaged.

The funnel: six stages, each with its own losses

“Too much information” hides six distinct stages, and Part II collapses if they blur — because each stage loses different things and each later mechanism owns a different transition:

AVAILABLE        everything published in reachable stores
      ↓
RETRIEVABLE      what queries and subscriptions can technically fetch
      ↓
SURFACED         what the system actually puts before the person
      ↓
NOTICED          what attention lands on
      ↓
CONSUMED         what is actually read, watched, heard
      ↓
RETAINED / ACTED what changes belief, skill, or action

Available exceeds retrievable (unindexed stores, paywalls, untranscribed audio). Retrievable exceeds surfaced (ranking cutoffs, subscription boundaries). Surfaced exceeds noticed (banner blindness, digest-skipping, the twentieth item in a list). Noticed exceeds consumed (opened-but-abandoned tabs — the graveyard of good intentions). Consumed exceeds retained (the paper read twice and still not usable when needed). A “better recommender” typically optimises one transition — retrievable-to-surfaced — while the losses stack at every other joint. The Radar of Chapter 20 must instrument the whole funnel, and each chapter now owns its transition: 12–13 move long sources toward dense units (consumed-per-minute), 14–15 filter by task and novelty (surfaced-per-goal), 16 rewires subscriptions (retrievable-per-interest), 17 ladders resolution (consumed-per-minute again, reversibly), 19 governs surfaced-to-noticed timing. This chapter builds the funnel; the Part fills it.

One Part I principle crosses the boundary intact, scaled up: false omission can be more expensive than extra consumption. Chapter 4’s false skip — the dismissed advisory that mattered — becomes Part II’s missed paper, missed contradiction, missed warning. Every selection policy in this Part is therefore scored with misses counted separately and weighted heavier, exactly as EXP-04 required. A filter that halves reading time while dropping the one decisive item has not saved attention; it has destroyed value at a discount. The asymmetry is structural to selection under scarcity, and nothing downstream may average it away.

The attention curve: the finding that pressures everything

EXP-11 below is kept extremely simple on purpose — no novelty modelling (Chapter 15 owns that), no relation-aware ranking (Chapter 16), no timing (Chapter 19). Fixed budget, frozen corpus, simple policies: chronological/subscription order, keyword filtering, embedding relevance, generic LLM relevance. The point is not to crown a policy but to establish the curve:

5 minutes  → X useful information
15 minutes → X + ...
30 minutes → diminishing return
60 minutes → ...

Two properties of that curve do the chapter’s real work. First, its shape: if useful information saturates early, the problem is ranking (the good items exist but hide); if it climbs steadily, the problem is coverage (attention itself is the bottleneck and compression must do more). The Part’s strategy differs completely between those worlds, so the curve is measured before the machinery is chosen. Second, its disagreement structure: independent adjudicators given the same budget and goal will select different “useful” sets. That disagreement is not noise to be averaged out — it bounds how much any selection policy can be “correct,” proves importance is not objective, and justifies everything personal in Chapters 14–16. A chapter that skipped this measurement could pretend selection has right answers; this one removes that comfort in advance.

A shared world: the Part II corpus commitment

Part I’s one noted gap was corpus fragmentation — each experiment its own toy dataset. Part II repairs it from the start, because selection, compression, novelty, and timing can only compose if they operate on the same world. The commitment, binding from EXP-11 forward:

Part II experiments deliberately begin as closed-world mechanism tests: the information universe is frozen so selection policies can be compared under controlled conditions. Success there does not establish performance in an open web or continuously changing information environment.

Later chapters may write “within the frozen information world” and inherit this boundary without re-arguing it.

Part II experiments increasingly operate on one frozen information world: a conference-scale media corpus (talks with video, transcripts, slides; surrounding papers, threads, advisories) with declared reader goals, human-adjudicated usefulness keys, and versioned freezes — extended, never replaced, as chapters need more.

EXP-11 freezes the core: the corpus, the goals, the keys, the budget protocol. Chapter 12’s segmentation, Chapter 13’s units, Chapter 16’s subscriptions, and Chapter 20’s Radar all build on that freeze, adding layers (events, units, profiles) without swapping the world. If a later chapter needs material outside the freeze, it extends the corpus with the same keying protocol and records the delta. Ten separate toy datasets would let each mechanism look good in its own world and fail in combination; one shared world forces composition from the beginning. The Radar is evaluated on ground its components were built on — as it should be, since personal deployment is also one continuous world, not ten fresh ones.

What this chapter earned

Selection is unavoidable once supply exceeds attention — structurally, not as a design preference. The funnel (available → retrievable → surfaced → noticed → consumed → retained/acted upon) locates every Part II mechanism at its transition. Value-per-attention frames scoring without a universal importance score. False omission inherits Chapter 4’s asymmetry at stream scale. The attention curve and its disagreement structure are the measured facts that justify everything conditional to come. Which items are important remains entirely open — deliberately, since answering it is the Part’s work.

The first source type is long video, where duration is a storage property rather than a semantic unit.

References

  • Simon, H.A. (1971). Designing Organizations for an Information-Rich World. In Greenberger (ed.), Computers, Communication, and the Public Interest, Johns Hopkins Press, pp. 37–52. Content-read (primary PDF): wealth-of-information/poverty-of-attention; attention-conservation via receiving/storing and filtering/transforming; “listen and think more than speak” as design principle. Used for the budget framing; modern overload rates NOT licensed.
  • Information-foraging theory (Pirolli & Card programme, Ch 04 base): reused for patch/scent vocabulary at stream scale; no new claims taken.
  • Recommender-systems literature: referenced only as engagement-contrast (optimising consumption vs user utility); specific results deferred to Ch 16’s evidence review.

Proposed experiment EXP-11: budgets, policies, and the curve

Status: PROPOSED. Corpus: the Part II freeze v0 — conference-scale media (video + transcripts + slides where available) plus surrounding papers/threads, with 3–5 declared reader goals and human-adjudicated usefulness keys per goal (independent adjudicators; disagreement recorded as data). Conditions: chronological/subscription order, keyword filtering, embedding relevance, generic LLM relevance — NO novelty, NO relation-aware ranking, NO timing variation (all owned downstream). Budgets: 5 / 15 / 30 / 60 minutes of consumption, enforced. Measures, reported separately: useful items seen, useful items missed (weighted heavier), irrelevant items consumed, time spent, coverage of adjudicated-important information, and the marginal-gains curve across budgets. Inter-rater disagreement on “useful” quantified as a bound on achievable selection correctness. Failure criteria: no diminishing returns observed (budget framing wrong for this corpus); adjudicators agree perfectly (importance objective after all — collapses the Part’s motivation, kept as a live risk); simplest baseline ties all rankers (ranking adds nothing; compression must carry the Part alone). Artifacts expected: frozen corpus + version manifest, goal pack, usefulness keys with disagreement records, budget protocol, attention curves per goal. What a positive result would not justify: which items are important in general, or anything about personal novelty — selection pressure established, selection policy not.