Chapter 64

What a Symbol System Would Have to Be

Chapter 56 showed the method and named the price. This chapter is the price list.

It is the most technical chapter in the Prestige, and it is technical on purpose. The Turn spent three hundred pages establishing that a Kantian conclusion with a ledger attached is a claim, and that the same conclusion without one is a posture. Having just conceded that the book’s grounding was analytic rather than earned, the least i can do is be specific about what earning it would involve.

The chapter has four movements. What a symbol system is, which turns out to be a prior question nobody in the grounding literature answers (Section 64.1). Where the machinery already in hand meets that account and where it falls short (Section 64.2). What a budget buys, which turns out to be the deepest of three instances of one move that Chapter 68 will gather (Section 64.3). And a construction — \(\RHO[\Bee]\), the calculus-making functor — which organizes several of the conditions attractively, lands on a category the Turn already built, and then fails, for a reason i will make as sharp as i can (Sections 64.4 to 64.6).

64.1 The prior question

The reference bootstrap problem asks us to construct a symbol system. It is therefore prior to ask what one is.

The literature is remarkably thin on this. Symbol grounding work generally treats symbol as understood and asks how to attach it to sensors. Representation learning treats discreteness as an architectural choice, to be imposed with a quantizer where convenient. Emergent communication treats a symbol as whatever the channel happens to carry. None of these supplies a criterion by which a proposed system could be judged inadequate, which is what one needs if the construction is to be a construction rather than a hope.

Goodman’s Languages of Art does supply one [141]. It offers a set of conditions distinguishing notational systems from other symbol systems, developed for the score / performance relation in music and dance and generalized from there. i adopt it provisionally — provisionally because it has known failure modes, taken up in Section 64.1.3, and because adopting it is a bet that its conditions are the right ones rather than merely the first ones available. The one prior attempt to put the theory to computational work is Goel’s, which derives from notationality a class of physical notational systems and argues that these are exactly the dynamical systems capable of realizing computational information processing [140]. That work predates everything in the current setting and has not been revisited for learned representations. So far as i can determine it is the whole of the prior art, which is itself a small scandal.

64.1.1 The conditions

Goodman distinguishes marks (physical inscriptions or utterances), characters (classes of marks that count as the same), and compliants (the things a character applies to, its compliance class). A score is a character. A performance is a compliant.

A notational scheme satisfies two syntactic conditions.

Condition 64.1 Syntactic disjointness

No mark belongs to more than one character.

Condition 64.2 Syntactic finite differentiation

For any two characters \(K \neq K'\) and any mark \(m\) not belonging to both, it is theoretically possible to determine that \(m\) does not belong to \(K\), or that it does not belong to \(K'\).

A notational system is a scheme satisfying three further semantic conditions.

Condition 64.3 Unambiguity

Each character has a single compliance class: not different compliants at different times or in different contexts.

Condition 64.4 Semantic disjointness

No two characters share a compliant.

Condition 64.5 Semantic finite differentiation

The analogue of Condition 64.2 on the compliance side.

The two finite-differentiation conditions are what exclude dense systems. A system whose characters are ordered so that between any two there is a third fails Condition 64.2, and this is Goodman’s account of the difference between a score and a sketch, or between the digital and the analogue.

64.1.2 The roundtrip is derived, not primitive

It is natural — and it was my own first reading — to take the criterion of adequacy to be losslessness of the roundtrip: from score to performance and back by transcription, and from performance to score and back, without loss of identity. That is indeed the property Goodman is after and it is the property that matters here. But it is a consequence of Conditions 64.164.5 rather than an additional stipulation, and the distinction is load-bearing. The roundtrip tells you when you have succeeded. The five conditions tell you what to check, and they decompose the requirement into parts that turn out to have very different relationships to the machinery this book has built.

Remark 64.1 What the framework does and does not do

Goodman’s conditions are conditions on a symbol system: a scheme together with a compliance relation. They presuppose the compliance relation as given and adjudicate whether it is notational. They therefore cannot select a compliance relation, and they are invariant under relabeling: if \(\Sigma\) is a notational system and \(\phi\) a bijection on characters, then \(\phi(\Sigma)\) is a notational system with the same compliance classes differently named. Goodman’s framework says what a solution to the reference bootstrap problem would look like. It cannot by itself produce one.

Remark 64.1 is why adopting Goodman is half the battle rather than progress on the bootstrap problem proper. It is also the formal seed of Observation 56.6: a framework invariant under relabeling cannot fix reference, and this one is invariant by construction.

64.1.3 Failure modes, and provisional repairs

The wrong-note paradox.

Goodman’s compliance relation is exact. A performance containing a single wrong note is not a performance of the work. Since performances without wrong notes are essentially never found, the theory as stated has no instances in musical practice. This is the standard objection and it is a real one. i record the shape of the repair now and return to it in Section 64.3: exactness should be relativized to a resource bound, so that two performances are the same performance when no affordable observation separates them. That is not an evasion, because in a cost-accounted setting “affordable” is a defined quantity rather than a gesture, and the move converts an absolute condition into a family of conditions indexed by budget.

Density is a property of the world, not only of the notation.

Finite differentiation can fail because the notation is badly designed, or because the phenomenon being notated is dense and no notation of it could succeed. Goodman does not clearly separate these, and for our purposes the separation is essential: we need to know whether a failure to notate is a failure of the agent or a fact about the environment. Question 64.2 states this as an open problem, and it is the one on the list i would most like answered.

The ontological commitments are separable.

Goodman’s theory of notation comes attached to strong claims about the identity of works of art. Recent work in aesthetics argues that the notational apparatus survives being detached from those commitments and reinterpreted epistemically [137]. i take that detachment for granted.

64.2 Where the machinery meets it, and where it does not

64.2.1 The two bisimulations

The natural translation of the roundtrip requirement into process-calculus vocabulary is a pair of bisimulations. Write \(\mathcal{S}\) for the space of scores, \(\mathcal{P}\) for the space of performances, \(\sem{-} : \mathcal{S} \to \pop(\mathcal{P})\) for the compliance map taking a score to its compliance class, and \(\tau : \mathcal{P} \to \mathcal{S}\) for transcription. The requirement is that \[\begin{align} p \bisim_{\mathcal{P}} p' &\iff \tau(p) \bisim_{\mathcal{S}} \tau(p') \label{eq:rt1}\\ s \bisim_{\mathcal{S}} s' &\iff \sem{s} = \sem{s'} \label{eq:rt2} \end{align}\] with \(\bisim_{\mathcal{P}}\) and \(\bisim_{\mathcal{S}}\) the respective behavioral equivalences. Read one way this says the two kernels agree; read another, that \(\tau\) and \(\sem{-}\) are mutually inverse on the quotients.

64.2.2 What bisimulation supplies

Conditions 64.1, 64.3 and 64.4 are quotient conditions and fall out of (64.1)(64.2) directly. Syntactic disjointness says the character map is a function on marks, which is what it means for \(\bisim_{\mathcal{S}}\) to be an equivalence. Unambiguity says the compliance class of a character does not vary, which is the statement that \(\sem{-}\) is well defined on \(\bisim_{\mathcal{S}}\)-classes. Semantic disjointness says compliance classes partition, which is (64.2) read as injectivity on the quotient.

This is a good fit, and it is why the translation is worth making at all. Bisimulation is the mathematically mature version of same behavior, it comes with coinductive proof methods, and in the reflective setting it is the equivalence the namespace logic already denotes.

64.2.3 Where Goodman adds

Conditions 64.2 and 64.5 are not quotient conditions, and bisimulation does not supply them.

Observation 64.1 Separation is not quotient

Finite differentiation is a separation condition: it requires not merely that distinct characters be distinct, but that their distinctness be effectively determinable from a mark. A bisimulation quotient can be perfectly well defined over a dense space, in which case every pair of distinct classes is distinguishable in the limit and no pair is distinguishable by any finite observation. Such a system satisfies (64.1)(64.2) and fails Condition 64.2.

That is the precise content of the claim that Goodman is compatible with the bisimulation approach and adds strictly more. What he adds is an effectivity requirement that process equivalence, being defined by a greatest fixed point, is systematically silent about.

And the point is not academic. It is exactly where the neural implementations fail, and they fail there without the vocabulary to say so.

LatPlan and the symbol stability problem.

LatPlan learns propositional and action symbols from unlabeled image pairs and plans in the resulting latent space [130]. Asai and Kajino subsequently identified what they named the symbol stability problem: the propositions produced by the state autoencoder behave stochastically, so the same input does not reliably produce the same proposition [131]. In Goodman’s vocabulary this is a failure of Condition 64.2 — one mark, no effective determination of which character it belongs to — discovered empirically, by people who ran into it as an obstacle to making a learned symbol system behave like a symbol system.

Vector quantization as a fix for one condition only.

The VQ-VAE bottleneck assigns continuous latents to a codebook by nearest neighbor [145]. This imposes syntactic disjointness and a separation margin by construction: read through Goodman, it is an engineering solution to Conditions 64.1 and 64.2 and to nothing else. It says nothing about the compliance side, which is why quantization alone does not produce symbols that mean anything. The subsequent literature’s efforts to make the codes “semantic” by additional training objectives are attempts to buy Conditions 64.364.5 that do not know that is what they are doing.

64.2.4 Why bisimulation metrics are the wrong repair

The reinforcement-learning literature has independently encountered the rigidity of exact bisimulation and responded with quantitative relaxations: bisimulation metrics [138] and their scalable descendants [139, 152, 136]. These are the correct response to the wrong-note paradox in that setting and the wrong response here.

Observation 64.2 The metric smooths what notation needs separated

A bisimulation metric replaces the equivalence with a distance and the quotient with a geometry. It thereby dissolves precisely the discreteness that Conditions 64.2 and 64.5 require. A representation shaped by a bisimulation metric is, in Goodman’s classification, closer to a sketch than to a score.

The tension is genuine and not merely terminological. One cannot have a symbol system and a smooth behavioral geometry at the same level of description. What one can have is a discrete system whose admissible discriminations are set by a resource bound.

64.3 What a budget buys

There is a move this book makes repeatedly and has not yet named. Something is put out of an agent’s reach; one asks whether the barrier is a matter of logic or of money; and the answer keeps coming back money. It has appeared twice already. Hennessy–Milner adequacy characterizes a bisimulation class positively, by formulae of the agent’s own logic, and says nothing whatever about whether the separating formula can be afforded — so what an agent cannot distinguish, it usually cannot distinguish for want of funds rather than for want of concepts. And Chapter 53 shows that a learner does not merely fail to see the distinctions it cannot afford; it fails to see that they are there, and reports a world with fewer kinds in it than the world has.

Here is the third instance, and it is the one that bites hardest, because it is about which marks are the same mark. Chapter 68 collects all three and argues that what they amount to is Kant’s thing-in-itself surviving translation into computation as an economic fact rather than a logical one. i mention that now so the reader can see where this is going; the argument for it is that chapter’s, not mine here.

The mortal-scientist framework already prices observation. Assays cost. Refutation is budget-relative. A verdict may be \(\bot\) because the budget ran out rather than because the world was silent. That is exactly the ingredient Observation 64.1 says is missing.

Conjecture 64.1 Finite differentiation as a theorem

In a cost-accounted setting with budget \(\sigma\), define \(x \bisim^{\sigma} y\) to hold when no assay affordable at \(\sigma\) separates \(x\) from \(y\). Then the quotient by \(\bisim^{\sigma}\) satisfies Conditions 64.2 and 64.5 with the budget as the separation constant, and the resulting family \(\{\Qs\}_{\sigma}\) is a filtered diagram of notational schemes refining as \(\sigma\) grows.

i state this as a conjecture rather than a proposition because two things are unverified. First, that \(\bisim^{\sigma}\) is transitive: affordability is not obviously composable, and a chain of pairwise-indistinguishable states may have distinguishable endpoints, which is the standard failure of relations of the form “indistinguishable by a bounded observer”. Second, that the cost-accounting monad interacts correctly with the extension of Section 64.4; see Question 64.5. If the first fails, the repair is to take the transitive closure and accept a coarser scheme, at the price of the separation constant no longer being \(\sigma\).

Should Conjecture 64.1 hold, the reading is worth having. A poorer agent has a coarser notation. Notational adequacy becomes a purchasable property rather than a binary one, the wrong-note paradox is resolved by making exactness relative to what can be afforded rather than by abandoning it, and the whole of Goodman’s digital / analogue distinction becomes a statement about budgets rather than about the metaphysics of inscription. That is the same move as the other two, applied one level further down, to the question of what counts as a symbol at all — and it has the same consequence, which is that a wall turns into a frontier and the frontier has a price on it.

64.4 Reflection is score construction, not performance

We now have a criterion and no way of satisfying it. What follows is an attempt to see how far the machinery of this book can be pushed toward satisfying it, carried as far as it will go and then stopped. It is offered as an analogy. Section 64.6 argues, at some length and against my own construction, why it is not more than that.

The reflective higher-order calculus takes names to be quoted processes [1]. For a process \(P\), \(\quo{P}\) is a name; for a name \(x\), \(*x\) is a process; and \(*\quo{P} = P\), \(\quo{*x} = x\). Quotation and dereference are mutually inverse.

It is tempting — and i made this mistake before correcting it — to read that immediately as the score / performance relation, with Goodman’s roundtrip holding by construction. It is not, and seeing why is what makes the rest of this section follow rather than intrude.

In pure \(\rhoc\), both sides of \(\quo{-}\) and \(*\) are syntax. \(P\) is a term; \(\quo{P}\) is a term-code; \(*\quo{P}\) is the same term again. Nothing here is a performance. What the operators do is build scores out of scores: quotation lets a score be embedded in another score as an atom, and dereference recovers it. The construction is recursive score construction, and its losslessness is the statement that the recursion is faithful — that no information is lost by descending a level and climbing back up. That is a genuine and useful property. It is also a property internal to the notation. A notational system all of whose compliants are further inscriptions is precisely the case Goodman analyses under the heading of scripts that denote other scripts, and it is not yet a system that says anything about a world.

Observation 64.3 Two readings of \(\quo{-}\) and \(*\)

In \(\rhoc = \RHO[\varnothing]\), \(\quo{-}\) and \(*\) are score-to-score: they witness the faithfulness of a recursion inside the notation, and there is no performance in the picture. The compliance relation is trivial, in the sense that every character complies with syntax.

64.4.1 Builtins turn the recursion into a correspondence

Now adjoin a builtin process \(b\) drawn from the host world. The syntax of quotation does not change: \(\quo{b}\) is formed exactly as \(\quo{P}\) is. Its semantics changes completely, because \(b\) is not a term of the calculus. It is a behavior in the world the computation is embedded in.

At that point, and only at that point, \(\quo{b}\) is a score and \(b\) is a performance in Goodman’s sense. \(\quo{b}\) is a static, inspectable, transmissible mark that can be handled by the calculus’s own machinery — sent on channels, pattern-matched, quantified over in the namespace logic — while the thing it codes is an event elsewhere, running under laws the calculus does not supply. Dereference \(*\quo{b}\) is then not a syntactic identity but a performance instruction: produce, in the host world, something indistinguishable from \(b\).

Observation 64.4 Where the correspondence enters

Adjoining builtins does not add a new operator. It re-types the existing one. The same \(\quo{-}\) that was recursive score construction over \(\rhoc\)-terms becomes, on the builtin sort, a score / performance correspondence; and the same \(*\) that was syntactic recovery becomes performance. Goodman’s roundtrip requirement therefore attaches to \(\rhoc\) at exactly the point where the calculus is extended by the world, and nowhere else.

This is why the extension is the interesting object and the pure calculus is not. In \(\RHO[\varnothing]\) there is nothing to ground and the roundtrip is a tautology — which is Observation 56.1 again, arrived at from the side of the notation rather than from the side of the ontology.

64.4.2 Two ways to adjoin the world

There are two ways to perform the extension, and Observation 64.4 explains why only one of them does the job.

Builtin names.

Adjoin name constants beyond \(\quo{\nil}\), understood as channels to the host world. This gives an interface but no theory of what is on the other end: the calculus can send and receive, and has nothing to say about the behavior it is interacting with. It is the applied-\(\pi\) move [129], where a term algebra of data is adjoined and the added material is inert with respect to the operational semantics. Crucially it does not re-type \(\quo{-}\): an adjoined name is a name, not the quotation of anything, so there is no performance for it to be the score of, and Observation 64.3 still applies. The roundtrip stays inside the notation.

Builtin processes.

Adjoin processes \(b\) drawn from the host world. Then \(\quo{b}\) is a symbol for \(b\), and \(*\quo{b}\) is a process in the host world indistinguishable from \(b\). This is the version that does the work, for the reason just given: it is the only one of the two under which quotation acquires a compliance class outside the syntax.

Remark 64.2 This is not psi-calculi

The psi-calculi framework parameterizes \(\pi\)-like calculi by nominal datatypes for terms, conditions and assertions, together with a channel-equivalence relation, a composition, a unit and an entailment relation, and subsumes applied \(\pi\), spi, fusion and concurrent constraint \(\pi\) as instances [134]. The adjoined material there is data: inert under the operational semantics, transmitted but not run. Adjoining builtin processes is a different move, because the adjoined objects have their own dynamics and the calculus ceases to be closed — part of the reduction relation is delegated to the host. Readers who know psi-calculi will otherwise assume this is an instance of it.

64.5 The calculus-making functor, and a category we already have

The natural way to organize this is to stop thinking of \(\rhoc\) as a calculus and start thinking of it as a functor from behaviors to calculi: \[\RHO[\Bee] ::= \nil \mid \Bee \mid \inn{y}{x}{P} \mid \out{x}{Q} \mid P \Par Q \mid *x, \qquad x, y ::= \quo{P}\] with the usual equations and rewrites, and with \(\Bee\)’s behavior specified in a source category \(\Cat\) and imported into a version of the parallel-composition rule. The pure calculus is the fiber over the empty behavior, \(\rhoc = \RHO[\varnothing]\), and everything the Turn did with \(\rhoc\) was done in that fiber.

Now the part i find genuinely satisfying, because it was not planned.

Chapter 10 built a category of theories, and built it carefully, because the obvious definition was wrong three separate ways. Objects are GSLTs; morphisms are pseudofunctors of context multicategories that preserve context-labeled bisimulation, with the bisimulation on the target computed only over the image of the source’s contexts (Definition 10.1). And to stop the resulting comparison being trivial, that chapter imposed two conditions on a morphism \(F : G \to H\):

Those are faithfulness and density, and they were introduced to make the expressiveness preorder on calculi mean something.

\(\RHO[-]\) is a functor into that category. Take the morphisms of \(\Cat\) to be behavior-preserving maps. There is then a unit \(\eta_{\Bee} : \Bee \to \RHO[\Bee]\) including builtins as processes, and:

Observation 64.5 Goodman’s adequacy is hosting and exhausting

Goodman’s two-way losslessness for the score / performance relation of Observation 64.4 is the statement that \(\eta_{\Bee}\) is hosting and exhausting in the sense of Chapter 10, with respect to behavioral equivalence.

i want to be careful about how much that is and how much it is not. It is not a solution to anything. What it is, is the observation that the pair of conditions this book introduced to compare one calculus with another turns out to be the criterion of notational adequacy when one of the two is a world instead of a calculus. The question the Turn asked was: how much of calculus \(G\) can calculus \(H\) host, and how much of \(H\) does \(G\) exhaust? The question here is: how much of world \(\Bee\) can \(\rhoc\) host, and does \(\rhoc\) mint anything \(\Bee\) cannot cash? Same two conditions, same category, one argument changed from a theory to a world. Goodman’s criterion of adequacy becomes a property one could try to prove, which is more than it was before.

And it became a little more provable while this was being written. i said above that hosting and exhausting “are faithfulness and density” as though that were a gloss. Remark 10.2 back in Chapter 10 makes it a statement: a morphism carries a map \(\Phi\) on contexts, and hosting and exhausting are faithfulness and density of \(\Phi\). If that is right, then Observation 64.5 is not an analogy between two pairs of conditions but a single statement about one functor: Goodman’s two-way losslessness is the demand that the map sending world-distinctions to notation-distinctions be faithful and dense. i do not prove it here, and i note that the world side of that functor is exactly the encoder this book does not build. But a criterion with a shape is easier to attack than a criterion with a description, and this one now has a shape.

64.5.1 Three further things the functorial reading suggests

Each is a suggestion rather than a theorem, and i flag which are checkable.

(i) The import into \(\Par\) has a known shape.

“Behavior specified in the source category, imported into the par rule” is the shape of a distributive law of syntax over behavior. Turi and Plotkin’s bialgebraic semantics takes operational rules given as a natural transformation \(\lambda : SF \Rightarrow FS^{\dagger}\), with \(S\) a syntax functor, \(F\) a behavior functor and \(S^{\dagger}\) the free monad on \(S\), and yields an operational model together with a canonical, internally fully abstract denotational model [86]. The relevant corollary is that for any such law, behavioral equivalence in the final coalgebra is a congruence on the operational model. If the import can be put in that form, bisimulation-as-congruence over \(\RHO[\Bee]\) comes free rather than being re-proved per \(\Bee\). If it cannot, congruence fails and every compositional result of the Turn stops transferring across the cut. This is checkable and unchecked, and it is the single most consequential unchecked thing in this chapter.

(ii) Unambiguity forces behaviors, not states.

If the objects of \(\Bee\) are states with dynamics, then \(\quo{b}\) codes a moving target: the same mark has different compliants at different times, which is exactly the failure of Condition 64.3. The repair is forced. The objects of \(\Bee\) must be behaviors — a subobject of the final \(F\)-coalgebra — so that quoting a builtin quotes its entire future rather than a snapshot. Then \(*\quo{b} \bisim b\) holds definitionally, and unambiguity holds because behaviors do not change over time even though states do. A consequence worth stating on its own: one cannot take \(\Bee\) to be the physical furniture of the world. \(\Bee\) must be the world already quotiented by an observational equivalence — which is where Conjecture 64.1 would enter, with \(\Bee = \Qs(W)\), and which ties the whole construction back to a budget.

(iii) The source category is constrained.

Name matching in \(\inn{y}{x}{P} \Par \out{x}{Q}\) requires deciding name equality, and names now include \(\quo{b}\). So structural congruence on \(\RHO[\Bee]\) is decidable only if equality on \(\Bee\)’s objects is at least budget-decidable. This is a restriction on which worlds can serve as source, and it is the formal descendant of the fact that builtin names are the base case of the well-founded coding by which quotation manufactures fresh names.

64.6 Why this is an analogy and not a solution

i now argue against the construction just presented, because the argument is short and decisive and it is better made here than by a reader.

  1. It consumes what it was meant to produce. \(\RHO[-]\) takes \(\Bee\) as an argument. The reference bootstrap problem is the problem of finding \(\Bee\): of individuating the world into behaviors worth having symbols for. The functor presupposes a solved instance of the problem it was introduced to address, and everything it then does is bookkeeping over that solution. This is the decisive objection and the rest are supporting. It is also, i note without pleasure, the same objection as Observation 56.1, one level of abstraction up: the first time we assumed the world was made of processes, and this time we assume it comes pre-carved into them.

  2. Functoriality entails relabeling invariance. \(\RHO[\Bee] \cong \RHO[\Bee']\) whenever \(\Bee \cong \Bee'\). So the construction cannot fix reference; it transports whatever reference \(\Bee\) came with. This is not an imported objection but a theorem of the construction, and it is the formal version of Remark 64.1 and the machinery behind Observation 56.6.

  3. Losslessness is stipulated where it is interesting. \(*\quo{b} \bisim b\) holds by construction, because we defined \(\quo{-}\) on \(\Bee\) to be exact. In the world, sensor-to-symbol transduction is lossy, noisy and partial, and the interesting question is what to do when the roundtrip fails. The construction has nothing to say about approximate compliance, which is where all the empirical difficulty lives.

  4. The par rule import is unverified. Section 64.5.1(i) is conditional. Until the import is exhibited as a distributive law of the required shape, congruence of bisimulation over \(\RHO[\Bee]\) is an assumption, and with it every compositional result that would be transported.

  5. Two interaction disciplines complicate the logic. Communication in \(\rhoc\) is name-mediated. If builtins react in the source category by something that is not a name rendezvous, then \(\RHO[\Bee]\)’s parallel composition hosts two kinds of redex, and the namespace logic’s modalities — which quantify over redex positions — require a sort distinction. The ladders computed in Chapters 22 and 23 would need recomputation, and it is not known whether their qualitative shape survives.

  6. Nothing here proposes candidates. Even granting all of the above, the construction is a specification language for grounded symbol systems, not a procedure for acquiring one. It says what the answer must look like. It does not search.

The honest summary is that the functorial reading is a good way to state several of Goodman’s conditions in a form where they might be proved, and no way at all to satisfy them.

64.7 Where a transformer belongs

The framing above assigns transformers a job, and it is not the job they are usually given in discussions of grounding. It is also not quite the job Chapter 21 gave them, and i have corrected that chapter rather than leave the two accounts standing side by side.

The slogan there was: the network is the transducer, the ecology is the epistemology. The second half is right. The first half is wrong, and it is wrong in an instructive direction. In the construction of Section 64.5 the transducer is \(\quo{-}\) and \(*\) extended to builtins. That map is exact and lossless by construction; a network cannot improve on it and is not needed for it. What is needed, and what nothing in the calculus supplies, is a proposal mechanism: given a stream of host-world observation, propose candidate builtin behaviors. That is objection (a) of Section 64.6 turned into a task, which is the most useful thing one can do with a decisive objection.

Two features make it a well-posed learning problem, which is more than can be said for “learn symbols from sensors”.

The success criterion is coverage, not reward.

\(\Bee\) is a presentation: a generating set such that the observable world lies in the image of \(\RHO[\Bee]\). This is a generators-and-relations problem with a definite answer, and minimality of \(\Bee\) is a statistical-complexity analogue rather than a hyperparameter to be tuned.

There is a principled candidate.

Causal states are already the coarsest coarse-graining of pasts retaining full predictive power [148], and by Section 56.8 a next-token-trained transformer represents belief states over them [147]. That makes causal states the natural first candidate for which distinctions deserve builtin names. The correspondence extends further than convenience: Crutchfield’s taxonomy of \(\epsilon\)-machines runs from finite-state through countable to uncountable and fractal, and that taxonomy is a density hierarchy — so Condition 64.2 is the condition separating its bottom rung from the rest, and the question “is this world notationally accessible to this agent at this budget” becomes “is the budget-quotiented \(\epsilon\)-machine of the host process finite”.

That last equivalence is, i think, the most promising single item in this chapter, because both sides of it are computable and neither side has been computed.

A second job for a network is amortized bisimulation checking. Assays are priced in the framework already, checking behavioral equivalence is expensive, and an approximate checker with calibrated error is a legitimate purchase if its errors can be priced. That is more speculative and i do not pursue it.

64.8 Open questions

These are the ones i would take money on being tractable.

Question 64.1 Transitivity of budget-relative indistinguishability

Is \(\bisim^{\sigma}\) of Conjecture 64.1 transitive? If not, what is the cost of taking transitive closure, and does the resulting scheme still separate at a rate related to \(\sigma\)?

Question 64.2 Density: agent or world?

Given a failure of Condition 64.2, is there a criterion separating “this agent cannot notate this world” from “this world admits no notation”? A candidate: the former is a statement about \(\Qs\) for the agent’s \(\sigma\), the latter about \(\lim_{\sigma \to \infty} \Qs\).

Question 64.3 Notationality of learned codebooks

Goodman’s five conditions are checkable properties of an encoder / decoder pair. What do VQ-VAE codebooks, BPE tokenisers and LatPlan’s state autoencoder score on them? “How notational is this tokeniser” appears to be a measurable quantity and, so far as i can determine, has never been measured.

Question 64.4 Is the import a distributive law?

Can the import of \(\Bee\)’s behavior into parallel composition be exhibited as \(\lambda : \RHO \circ F \Rightarrow F \circ \RHO^{\dagger}\)? If so, congruence is free by [86]. If not, what weaker condition holds, and does the lax version suffice?

Question 64.5 Does \(\RHO\) preserve the cost-accounting monad?

The Turn works inside plain \(\rhoc\) and appeals to a WLOG on syntax that is valid there. Do \(\RHO[-]\) and the cost monad of Chapter 11 commute? If they do not, the foraging inequality, the pricing schedule and the monotonicity of cost in distance do not transport to \(\RHO[\Bee]\), and with them Conjecture 64.1 loses its setting.

Question 64.6 The minimal social configuration

Observation 56.6 says grounding is a property of a composite. What is the smallest composite that suffices? Two agents with shared action, or does the argument require a population? Does the typed-overlap apparatus of Chapter 23 already contain the answer?

Question 64.7 Approximate compliance

Section 64.6(c) observes that the interesting case is roundtrip failure. Is there a notion of graded compliance that degrades Goodman’s conditions gracefully without collapsing into the metric of Observation 64.2?

64.9 What is claimed and what is not

Claimed.

That the correspondence between a language and a world is a fact about the language’s users and is not contained in the corpus (Observation 56.2); that the record of decipherment and of cetacean bioacoustics is evidence for this rather than merely consistent with it, and that the two successful routes both acquire the correspondence from agents who already have it (Observation 56.3); that next-token prediction recovers the structure of the token-generating process, which for a corpus is the linguistic community and not the environment (Observation 56.4); that the problem has at least two independent solutions in nature, so that solubility is not in question and impossibility arguments are unsound (Observation 56.5); that this book’s grounding is analytic and therefore does not address the problem (Observation 56.1); that Goodman’s conditions decompose into quotient conditions bisimulation supplies and separation conditions it does not (Observation 64.1); that bisimulation metrics are the wrong repair for the latter (Observation 64.2); that quotation and dereference in the pure calculus are score-internal recursion and acquire a compliance class outside the syntax only when builtin processes are adjoined (Observations 64.3 and 64.4); that Goodman’s adequacy criterion is hosting-and-exhausting for the unit of \(\RHO[-]\) (Observation 64.5); and that no relabeling-invariant single-agent framework can deliver correspondence to our symbols (Observation 56.6).

Spent immediately.

The hosting-and-exhausting reading of Goodman is used in the very next chapter and it is worth saying what for. Chapter 65 observes that the pair of conditions is exactly what separates the versions of the simulation hypothesis from one another: a simulation that never leaves the level it simulates admits an encoding that is both hosting and exhausting, and is therefore inert; one run from a higher position in the lattice cannot be exhausting, and what its inhabitants bootstrap is a shadow. i had not expected the criterion to have that use, and i note it here because a criterion that turns out to answer a question it was not built for is worth more than one that does not.

Conjectured.

That cost accounting supplies finite differentiation with the budget as separation constant (Conjecture 64.1).

Offered as analogy only.

The entire content of Sections 64.464.5. The functorial reading organizes the conditions attractively and suggests several checkable statements. It does not bootstrap anything, for the reasons in Section 64.6, and the first of those reasons is not a technicality.

Not claimed.

That Goodman’s conditions are the right ones — they are adopted provisionally, and the density and wrong-note objections are live. That transformers cannot in principle transduce — only that next-token prediction over an already-grounded corpus is not transduction, and that the two are routinely conflated. And that the reference bootstrap problem is soluble by construction: nature’s two solutions were achieved by populations, under selection, over long timescales, with interfaces that co-evolved with the symbol systems they support, and this book has no argument that any of those conditions can be dispensed with.