Related Work, and Where This Differs

Related Work{Related Work} {table}table{0}

A related-work section normally goes at the back, where it can be read by people who already agree with you. This one is at the front because the book is long, and because a reader is entitled to know before investing which existing programme this most resembles and why it is not that one.

There is a line to carry through all of it. Of every programme discussed below, this is the only one in which the knower’s epistemic limits are priced. Not bounded, not idealised away, not left to an implementation — priced, in a conserved resource, on a ledger the agent cannot refill from outside. Everything else i might have offered as the differentiator turns out to be shared with somebody. That one is not, and the rest of this chapter is partly an attempt to check that claim against the strongest neighbours i can find.

For each neighbour i try to say four things: what it claims, what this book takes from it, where they part company, and whether the parting could be checked by anything. The last is the one that makes this more than a bibliography with opinions.

Theories of mind and consciousness

Integrated Information Theory.

Tononi’s \(\Phi\) is a graded measure of consciousness computed from a system’s causal structure [18], and this book also grades rather than thresholds — a point Chapter 63.2 used to make in a sentence and which deserves more. The parting is over what the grading is a measure of. \(\Phi\) is intrinsic: it is a property of a system’s own cause-effect structure, computed by an analyst standing outside it. The coordinate here is relational: it is a position in a lattice of choice strength, and it is a fact about what an agent can do at a cut with another agent on the far side. The two come apart on a case that is easy to state. A system with rich internal causal structure and no interaction surface — nothing it can meet, nothing it can be met by — has, on IIT’s account, whatever \(\Phi\) its structure gives it. On this account it has no coordinate at all, because a coordinate is a fact about apertures and it has none. Whether that is a reductio of one view or the other is exactly the question, and it is at least a question with an answer. A second and less flattering difference: IIT has an experimental programme and this does not.

The free energy principle.

Friston’s account [17] is the closest neighbour in the book, and the resemblance is real: both make prediction the central activity of a persisting system, and both make the failure of prediction the thing that threatens persistence. The parting is that the free energy principle has no ledger. Prediction error is minimised, but nothing is spent to minimise it, and there is no account that can run dry. That difference is not cosmetic, and Chapter 21 shows where it bites: once assays cost, being right stops being the same thing as being fed, and a cruder theory that costs a fraction as much can dominate the true one. An agent minimising prediction error alone would not do that. An agent on a budget provably will.

Affect as homeostasis.

Solms [14] is the free energy principle pushed onto feeling: affect is the registration of departure from homeostasis, and consciousness is that registration extended. It is the direct ancestor of Chapter 26, which is the same argument rebuilt with a ledger where the biology was, and the debt is acknowledged at length in that chapter rather than here. The parting has two parts. The first is the one just made: Solms inherits the missing account, so nothing in his story can run out, and the price of internalization has nowhere to sit. The second is a disagreement about reach. Solms means the homeostatic argument to arrive at consciousness; this book takes it to arrive at sentience and stops it there, on the ground that occupying a position in the lattice is a different structure from owning a state, and that Chapter 62 needs machinery Chapter 26 neither has nor wants. One asymmetry should be said plainly, since it runs the other way from most of this section: Solms has clinical and lesion evidence and this book has none, and the concession made below about the missing experimental arm is not repaired by acquiring a neighbour who has one.

The honest concession used to be in the other direction, and this paragraph is the record of it having moved. In earlier drafts \(\varsigma\) was a distinguished component of the learner’s account driven by an externally supplied reward, which made it a budget line rather than a feeling: one drive, no valence, no arousal, nothing that was about anything. Chapter 26 withdraws that and rebuilds it as a formula the learner holds about its own reserves, so \(\varsigma\) is now signed, is a vector indexed by token flavor, and is about the reservoir it reads. The reward function stops being a free parameter, because the punishment is the learner’s own assay reporting that its goal is receding, and value conflict acquires a definition — two deficits with no arbitrage-free conversion between them, which is Chapter 24’s incommensurability wearing a different hat.

What remains conceded is smaller but real. There is still no arousal dimension, the setpoint’s placement is inherited rather than derived, and whether internalization pays is a conjecture rather than a result. This is a theory of affect now, in the sense that it has the right moving parts. It is not a good one yet.

Global workspace.

Baars’s architecture [153] and Dehaene’s neuronal implementation [157] make consciousness a matter of broadcast: content becomes conscious by winning access to a shared channel that many specialised processes can read. Chapter 23 tells a broadcast-and-contention story of its own, and the resemblance is close enough to be worth naming. The parting is over where the architecture comes from. Global workspace theory posits the workspace, because brains appear to have one. Here the composite is forced: a single learner revising its theory incrementally provably cannot leave the region it started in, so the unit of learning has to become a population sharing a namespace, and the shared namespace is the workspace arriving as a consequence rather than as a design.

Attention schema.

Graziano’s proposal [161] is that awareness is the brain’s model of its own attention — a self-model, and a deliberately imperfect one. This is the neighbor whose absence i felt most, and the absence is now partly filled. Chapter 26 is a learner modelling one thing about itself — its own reserves — and paying the ordinary tariff to do it, which is the shape Graziano’s proposal has. Worse, the machinery for the rest is already present: Section 68.5.1’s argument, applied reflexively, gives an agent that cannot afford to resolve its own structure and therefore reports a simplified one — which is the introspective illusion, derived rather than assumed, in about three sentences. Those three sentences are still not in this book, and they should be. What Chapter 26 supplies is the account a learner models, not the attention it models; Remark 26.6 gets as far as constraining the space of admissible attention policies and declines to pick one.

Chalmers, and the hard problem.

The Pledge argues that the productive question is not what it is like to be a bat but what it is like to be a computer program [166, 154]. It should be said plainly what that does and does not accomplish. It does not dissolve the hard problem. It relocates it into a setting where the surrounding questions become tractable — what an agent can represent, what it can afford to represent, what it must be committed to, what it cannot distinguish and cannot tell that it cannot distinguish — and leaves the residue exactly where Chalmers left it. Making a question tractable is worth doing. It is not the same as answering it, and this book does rather more of the first than the second.

Computational and physical foundations

Wolfram’s physics project and the ruliad.

The nearest neighbour in ambition [171, 172]: space, time, and a good deal of physics derived from rewriting rather than presupposed. Two partings, and both are structural. First, a GSLT carries equations and a named site of interaction, so the cut at which two terms meet is part of the presentation rather than something that emerges from the rewriting; that is what makes it possible to meter anything at all. Second, and more sharply: this book fibres over a category of theories, in which no object is privileged. The definite article in “the ruliad” is doing work that this framework declines to do. The rule space here is a variety of possible physics (Part Part IV), and which one you get depends on which further axioms hold — a difference that shows up immediately in what each programme thinks it is explaining when it recovers a piece of physics.

Constructor theory.

Deutsch and Marletto’s programme [158, 164] is the nearest neighbour in method, and the resemblance is striking once seen. Both replace statements about how a system evolves with statements about which transformations are possible. Both take that reformulation to be more fundamental than the dynamical laws it subsumes. And the parting is unusually clean. Constructor theory’s possible/impossible is Boolean. Here it is priced. That is not an analogy: it is exactly the distinction Chapter 30 already draws between decorating a theory in the Boolean semiring and decorating it in the tropical one. Constructor theory is, on this reading, the Boolean fibre of a construction this book runs over a family of semirings, and possible-at-a-cost is the tropical one. If that identification survives being made carefully, it is the most useful single sentence in this chapter, and i do not think it has been made before.

It from bit, and the computational universe.

Wheeler [170] and Lloyd [163] are ancestors rather than rivals, and the debt is general. What this book adds to the slogan is a specific answer to which computation, argued for in Section 4.2 rather than assumed: not a function machine, because the world is interactive, and not merely an interactive one, because an agent has to be able to represent its own hypotheses.

Applied category theory.

The categorical machinery here is standard and the debts are to a tradition rather than to a result: Spivak and Fong’s presentation of the field [159], Coecke and Kissinger’s diagrammatic method [155]. One honest gap belongs in this paragraph. The cost monad and the history monad are applied together throughout the book, and no distributive law between them is ever stated. Remark 35.2, on the two erasures, is very close to being the observation that they interchange only laxly, and it stops short of saying so.

The technical lineage.

Milner’s process calculi [84], Girard’s linear logic, Abramsky’s proofs-as-processes [99], Turi and Plotkin’s bialgebraic semantics [86], Turing’s oracle hierarchy [12], and the Weihrauch degrees [79] are the actual apparatus, cited throughout and not argued with. The one place the book takes a position against the tradition is Chapter 61, which claims that linear logic types the communication skeleton of a computation and is silent by construction about how its nondeterminism resolves — so that a logic of choice is the complement of linear logic rather than a refinement of it.

Learning, agency, and the symbol problem

Solomonoff induction and AIXI.

Solomonoff’s prior [168] and Hutter’s AIXI [162] are the canonical account of optimal prediction and optimal action given unbounded computation. The mortal scientist is the same problem with a budget and a body, and the relationship is sharper than a slogan. Proposition 21.3 says the complete theory of an environment sits at the top of the lattice as a limit no learner with a finite token supply can install — which is to say that AIXI’s optimality is precisely the thing this framework proves to be unpurchasable. The two accounts are not competitors so much as the two ends of one axis, and the interesting content is entirely in what happens between them.

Computational mechanics.

Crutchfield and Shalizi’s \(\epsilon\)-machines [148] are load-bearing in the Prestige rather than merely adjacent. Causal states are the coarsest coarse-graining of pasts that retains full predictive power, and the proposal of Chapter 64 is that they are the natural first candidate for which distinctions in a host world deserve builtin names. The taxonomy of \(\epsilon\)-machines — finite, countable, uncountable, fractal — turns out to be a density hierarchy in Goodman’s sense, which converts a question about notational adequacy into a question about whether a budget-quotiented \(\epsilon\)-machine is finite. Both sides of that equivalence are computable. Neither has been computed.

Symbol grounding and emergent communication.

Harnad’s problem [142], Goodman’s theory of notationality [141], Steels’s language games [150], and Taniguchi’s collective predictive coding [151] are treated at length in Chapters 56 and 64 and are only noted here. The short version: the correspondence between a language and a world is a fact about the language’s users, the emergent communication literature supplies the population that a single-agent account structurally cannot, and neither solves the problem this book concedes it has not solved.

Bisimulation metrics.

The reinforcement learning literature has independently hit the rigidity of exact bisimulation and answered with quantitative relaxations [138, 139, 152, 136]. Observation 64.2 argues this is the right repair in that setting and the wrong one here, because a metric dissolves exactly the discreteness a notation requires. This is a genuine disagreement with active work and i would rather state it as one than leave it in a footnote.

Hyperon and MeTTa.

Goertzel’s programme [160] is the closest thing to a sibling project, and its near-absence from this book is the most conspicuous omission in this chapter. Both take a rewriting substrate to be the right foundation for general intelligence; both make reflection central; and MeTTaIL, the language-definition platform much of the machinery here was implemented against, sits directly between them. The comparison i owe and have not made is between MeTTa’s approach to self-modification and this book’s — reproduction as recombination on quoted code — and between Hyperon’s attention allocation and the metering here. That is a chapter, not a paragraph, and it is not written.

Life and origins

Assembly theory.

Cronin and Walker’s programme [156, 169] is the largest single engagement in the book. What is taken: the assembly index as a measure of causal depth — the minimum number of steps by which an object can be built with reuse — and the claim that depth combined with abundance is a biosignature. What is offered back: answers to two questions the theory leaves open, namely what counts as a copy and how to specify the scope within which copies are counted, both of which Chapter 53 treats as fixed-point problems with computable answers. Where it dissents: Proposition 53.3 says that a coarse instrument does not merely mismeasure copy number, it manufactures copies, at a rate exponential in the coarseness, and the manufactured copies are biased toward the biosignature the measurement is looking for. That is a criticism of the measurement protocol rather than of the theory, and it should be read as such. It is also, so far as i know, the only claim in this book that a laboratory could check next year.

Smith and Morowitz.

The phase-transition account of the origin of life [52] supplies the picture of a thermodynamically driven whole within which chemistry is channelled. Chapter 60 reaches the conclusion that Smith’s thermodynamic macro-component and the determinacy islands of this framework cross rather than nest — and an earlier reading of my own, which had them nesting, is retained in a remark explaining why it does not survive.

Autopoiesis and \((M,R)\) systems.

Maturana and Varela [165] and Rosen [167] are the prior attempts at organisational closure, and Rosen is the one that should have been in this book from the start. “Closed to efficient causation” — the requirement that every function of an organism be produced by the organism — is close kin to the replication-as-a-fixed-point construction of Chapter 49, close enough that the difference is worth stating. Rosen’s closure is a condition on a category of mappings, and his conclusion is that such systems are not simulable by machines. The closure here is a fixed point of a reflective rewrite theory, and the same structure is exhibited constructively. Whichever of us is wrong, we are wrong about something specific, which is more than most pairs of positions in this area manage.

The table

Table 0.1 scores the programmes on five questions. The last column is where the argument of this chapter lives, and it is a claim rather than a description — which is why the column exists and why it is last.

{1.35}

ProgrammeTheory ofPrimitiveObserver inside?What would falsify itAnything priced?
IITconsciousnesscause–effect structurenoa high-\(\Phi\) system reporting nothingno
Free energypersistenceprediction erroryespersistence without error minimisationno
Global workspaceaccess consciousnessbroadcastpartlyconscious report without broadcastno
Constructor theoryphysical lawpossible / impossible transformationsnoa task both possible and impossibleboolean only
Ruliadphysicsrewritingyesa physics not in the rule spaceno
AIXIoptimal agencyalgorithmic probabilityno— (optimal by construction)no
Assembly theorylife detectioncausal depth \(\times\) abundancenohigh assembly index abioticallydepth, not access
This bookall of the above, at a costmetered interaction at a cutyesfinite contention; the copy-manufacture boundyes
Table 0.1 Five questions, scored. The final column is the claim this chapter is making: the framework developed here is the only one on the list in which what an agent can know is limited by what it can afford rather than by a prohibition. Constructor theory scores partially because possible/impossible is a pricing with two prices. Assembly theory scores partially because it prices construction and says nothing about the cost of access.

Two rows deserve comment because they are the ones i expect argument about.

Constructor theory is marked “Boolean only” rather than “no”, and that is deliberate: a two-valued possibility measure is still a measure, and on the reading proposed above it is precisely the Boolean fibre of the semiring-parametric construction of Chapter 30. Whether that is a generalisation of constructor theory or a misreading of it is a fair question and i do not think the answer is obvious.

AIXI’s falsification cell is a dash, and this is not a cheap shot. AIXI is optimal by construction relative to its own prior, so there is no observation that could refute it; what can be refuted is the claim that it is a useful model of anything realisable, and Proposition 21.3 is one form of that refutation. A framework that cannot be wrong is doing something other than what this book is trying to do, and the dash records that rather than scoring it.

Two concessions

Where this is a special case rather than a rival.

The account of interaction, the modal logic, and the categorical apparatus are all standard, and nothing in Part Part I is offered as a departure from the process-calculus tradition. The novelty, if there is any, is in what happens when a meter is installed and an agent is made of the same material as its environment. A reader who wants to describe the whole of the Machinery as “known, assembled unusually” will get no argument here.

Where a neighbour is ahead.

Two of these programmes have empirical arms and this one does not. The free energy principle has a substantial experimental literature. IIT has a perturbational complexity index and clinical data behind it. Assembly theory has a mass spectrometer and a published protocol. This book has falsification conditions — finite contention in the massive-population argument, and the copy-manufacture bound of Chapter 53 — and they are Fermi estimates rather than experiments. The second of the two is the closer to being run, and it is the thing i would most like someone to take away from this chapter.

{.table}table{0}