Chapter 26
Suffering: A Learner’s View of Its Own Account
Everything the ecology has done so far, it has done with persistence.
A learner is funded or it starves. It forages, it converts, it pays for its assays, and if the account will not cover the next rewrite the next rewrite does not happen. Selection operates on that directly: the lineages one finds are the lineages whose members kept going. Nothing in this account requires that a learner know any of it. The bookkeeping is ours. The learner has an account in the way a stone has a mass — as a fact about it, not as a fact available to it.
This chapter asks what changes when the learner gets to look.
The word is overloaded in this book and the collision is unfortunate enough to be worth heading off.
Throughout Part Part I and Chapter 24,
persistence is a syntactic property of a receive: the
<= form, a process that continues to offer after it has fired,
as against the linear form that consumes itself.
That is the sense in Definition 24.2 and in
Remark 24.3.
The sense meant in the opening paragraph above, and nowhere else in this chapter, is the selectional one: a lineage persists if its members go on existing. Remark 21.17 already had to separate these — a term that reinstates itself and a term that spawns a copy look alike until the ledger distinguishes them — and the reader who keeps that remark in mind will not be caught. Where confusion is possible below we say mere survival and let the technical word alone.
26.1 The move: an account a learner can quote
The apparatus for looking is already built, in three pieces that have not previously been put together.
The first is the account itself. Chapter 14 replaced the single phlogiston number with a vector: a learner carries balances indexed by signature, so its state is a pair \(\langle P, \vect{A}\rangle\) with \(\vect{A} = (A_1,\ldots,A_k)\), and different rules debit different components. Part Part IV will state the discipline formally, as Definition 30.6, and will then spend a chapter repairing a defect in it — that the definition says nothing about who may draw on \(\vect{A}\). Nothing below turns on the repair. What this chapter needs is only that the account exists, that it has components, and that a rewrite happens when the account covers what the rule charges. By Chapter 24 the resource types with which an ecology actually deals are flavors: signature-indexed tokens which are not freely interconvertible, whose exchange rates are set by enzymes, and whose residual incommensurability is the subject of §24.10. So the thing to be watched is a vector, not a scalar. This will matter more than anything else in the chapter.
The second is reflection. By Definition 21.12 a learner can hold \(@P\), a name for its own code, and by §21.3 a name predicate sees into a quotation — at reflection prices, which by Table 21.1 are the cheapest access the calculus sells. Nothing in that argument was specific to code. A learner that can quote its own program can quote the located reserves that fund it.
The third is instruments. Chapter 20 showed how to make spatial structure visible to bisimulation by placing a probe on one side of the cut, so that what was a structural inspection becomes an interaction and is therefore subject to the ordinary discipline of cost. An account is spatial structure of exactly the kind that construction was built for.
A learner \(\Lrn\) is internalizing if it holds a name \(a_\Lrn = @\vect{A}\) for its own account and an admissible instrument \(\Ob\) in the sense of Definition 20.3 whose probe reads \(a_\Lrn\), so that formulae of \(\HML(\ctx)\) about \(\vect{A}\) are among the hypotheses \(\Lrn\) can assay. A learner which is not internalizing is merely surviving.
No primitive is added. The account was already a located structure, quotation was already available, and the instrument discipline of Chapter 20 was already the price list for turning structure into observation. Definition 26.1 only says that a learner may aim that apparatus inward, and stipulates that when it does, it pays the same tariff as when it aims it outward.
That last clause is the whole of the chapter’s honesty. A framework in which introspection were free would prove anything one liked about it.
26.2 Homeostasis as a held formula
An internalizing learner can state hypotheses about its own reserves. The interesting thing it can do with that capacity is adopt a goal.
A setpoint for \(\Lrn\) is a vector \(\setpt \in A^k\) together with a tolerance \(\varepsilon \in A^k\). The associated comfort region is \(\mathcal{H} = \{\, \vect{A} : |A_i - A^\star_i| \leq \varepsilon_i \text{ for all } i \,\}\), and the homeostatic formula \(\phi_{\mathrm{home}} \in \HML(\ctx)\) is a formula whose extension over accounts is \(\mathcal{H}\).
The bathtub is the right picture and it is worth spending a sentence on, because the picture carries the content. A tub has an inflow it does not set and a drain it partly sets, and a level which is the running integral of the difference. Holding the level is not achieved by any single act; it is achieved by matching one rate to another, continuously, with the level itself as the only evidence that the matching is working.
Let \(\vect{A}(t)\) be the account of \(\Lrn\) at cycle \(t\). The drift is the signed vector \(\drift(t) = \vect{A}(t) - \setpt\), and the drift rate is \(\drift(t) - \drift(t-1)\), which is the componentwise difference of inflow and outflow over the cycle.
Nothing so far is a feeling. Drift is a fact about the learner, and up to this point it has been the kind of fact the learner does not have.
26.3 Suffering
The suffering of an internalizing learner \(\Lrn\) at cycle \(t\) is the assayed drift: the value \(\suff(t)\) that \(\Lrn\)’s own instrument returns when it evaluates \(\phi_{\mathrm{home}}\) and its graded refinements against \(a_\Lrn\). It is signed, it is vectorial, and it is indexed by flavor: \(\suff(t) = (\suff_1(t),\ldots,\suff_k(t))\), with \(\suff_i\) the learner’s own reading of how far and in which direction \(A_i\) stands from \(A^\star_i\).
Three things about that definition matter, and each of them is a correction to the account this chapter replaces.
Suffering is not a component of the account. It is a formula about the account. Earlier drafts of this book carried a distinguished component \(A_{\mathrm{s}}\) of \(\vect{A}\), singled out by fiat and driven by an externally supplied reward. That object is withdrawn. It confused a quantity the learner has with a quantity the learner reads, and the difference between those two is the entire subject of this chapter.
The punishment is not administered. The construction it replaces posited reward and penalty magnitudes \(\delta^+\) and \(\delta^-\), applied to the learner from outside whenever a prediction resolved. Where those numbers came from was never said, and they were the weakest joint in the argument. Here there is nothing to supply.
Let \(\Lrn\) be internalizing with setpoint \(\setpt\). Then the sign and magnitude of \(\suff\) are determined by \(\vect{A}\), \(\setpt\) and \(\Lrn\)’s instrument, with no further parameter. In particular no reward function is a free parameter of the framework: what plays the role of punishment is \(\Lrn\)’s own assay reporting that \(\phi_{\mathrm{home}}\) has stopped holding.
Immediate from Definitions 26.2, 26.3 and 26.4: each is a function of data already fixed. The only choice remaining is the choice of \(\setpt\) and \(\varepsilon\), and §26.5 argues that this choice is not the learner’s either.
The word “punished” should be read with care, because the framework does not contain a punisher. The learner holds a goal; the learner reads its own state; the reading says the state is receding from the goal. That is the whole mechanism. There is no valuation applied from outside and no bookkeeper marking it down. It suffers because it looks, and because it holds a standard against which looking can go badly.
A learner that held no setpoint could not be punished in this sense, and would be no worse off for it in any respect the previous chapters could measure. That is precisely why the previous chapters got by without this one.
Suffering is causally efficacious, but not by being money. The old account made feeling matter by making it a budget line: an agent that felt bad had less to spend. That conflated a mood with a liquidity crisis, and it made the theory unable to distinguish suffering from poverty. The present account separates them.
There exist internalizing learners \(\Lrn_1, \Lrn_2\) over the same resource algebra with \(\vect{A}(\Lrn_1) \geq \vect{A}(\Lrn_2)\) componentwise and \(|\suff(\Lrn_1)| > |\suff(\Lrn_2)|\).
Take \(\Lrn_2\) with \(\setpt_2 = \vect{A}(\Lrn_2)\), so \(\suff(\Lrn_2) = 0\): it is exactly where it means to be. Take \(\Lrn_1\) with \(\vect{A}(\Lrn_1)\) larger componentwise but \(\setpt_1\) larger still. Both readings are well defined and the inequality holds.
Proposition 26.2 is trivial to prove and was impossible to state under the old definition, which is the point. A rich learner far from its setpoint suffers. A poor learner at its setpoint does not. Anything that hopes to be a theory of affect rather than a theory of budgets has to be able to say that, and a distinguished account component never could.
26.4 Why this is affect and not a budget line
It is worth auditing the construction against what one would want of anything called a feeling, since the book has been criticized on exactly this point and the criticism was fair. Before the audit, an acknowledgement, because the criticism is not the only thing this section is answering to.
Nothing above was arrived at independently, and it would be graceless to let the reader think otherwise.
Mark Solms argues in The Hidden Spring [14] that affect is not a decoration on cognition but the registration of departure from homeostasis: a self-organizing system must hold a narrow band of internal states against a world indifferent to whether it does, feeling is how distance from that band presents itself, and this — rather than perception, or report, or any cognitive achievement — is what wants explaining first. This chapter is an attempt to find what survives of that argument when the biology is taken out and nothing is left in its place but a ledger. Definition 26.2 is his narrow band, written as a formula a learner can hold about its own account; Definition 26.4 is his registration, written as an assay; and Proposition 26.1 is why the exercise seemed worth the trouble, since it reports that once the band and the assay sit inside the learner rather than inside our description of it, the reinforcement signal stops being something we get to choose.
Solms builds on Friston’s free energy principle [17], and inherits from it the gap this book parts from everywhere: free energy is minimized, but nothing is spent minimizing it, and no account can run dry. Section 26.5 is what changes when one can. The price of internalization, and the fact that no individual is in a position to decide to pay it, are not in Solms and are not in Friston. They are what the ledger adds.
A second debt is less comfortable and had better be stated in full. Solms’s case is empirical: children born without a cortex and decorticate mammals remain affectively responsive, while small lesions of the upper brainstem abolish consciousness where large cortical lesions do not. This book has no empirical arm and does not acquire one by citation. What that evidence supports is that in the animals we actually have, affect is homeostatic registration — which is to say that the shape of the argument above is the shape of something selection built. That is convergence and not confirmation, and it is worth having on those terms and no better ones. Conjecture 26.1 is where the two could be made to meet, since a claim that internalization is selected by environmental variance is a claim comparative biology is equipped to embarrass.
The accounts part on what the argument is an argument about. Solms takes affect to be consciousness; the homeostatic story is meant to reach all the way, and one of his central claims is that consciousness is an extended form of homeostasis. Here it reaches sentience and is stopped there. What Definition 26.4 yields is a learner with a state of its own that its wealth does not fix — something it owns. It is not a learner that inhabits a different world from its neighbors, and Chapter 62 argues that the second is a different structure carrying a different price: a position is occupied, not owned, and occupying one is not something the ecology of Part Part II can pay for at all. Solms’s remaining claim, that the seat of the thing is the upper brainstem, is the most contested part of his case, and nothing here needs it. The two claims this chapter borrows are the two that do not turn on the anatomy.
One difference of vocabulary, since it is the sort of thing that causes arguments. Solms says affect, and unpleasure, and is careful about it. We say suffering, which is worse behaved, and §26.6 says why we are keeping it.
It is multi-dimensional. \(\suff\) is a vector indexed by flavor, and by Chapter 24 the flavors are not freely interconvertible. Two learners with the same aggregate shortfall may have it in different components and behave differently in consequence. Valence is the sign; magnitude in a component is that component’s urgency.
It is about something. \(\suff_i\) is indexed by a flavor, and a flavor is a signature, and a signature names a community. A learner short in the flavor it obtains by predation is in a different state from one short in the flavor it obtains from a source, and the index says which. This is aboutness of the only kind the framework can supply — a namespace — but it is not nothing, and it was entirely absent before.
It modulates policy and not only budget. This is the substantive claim of the section and it needs the setpoint to make sense.
Let \(\Lrn\) be internalizing and let \(\sigma\) be an allocation of its metered search budget across the branches of Proposition 21.6. Then two learners with identical \(\vect{A}\) and identical hypothesis stacks, differing only in \(\setpt\), admit different optimal \(\sigma\).
The value of an assay to an internalizing learner has two parts: what it tells the learner about the world, and what it tells the learner about \(\phi_{\mathrm{home}}\). The second part depends on \(\drift\), hence on \(\setpt\). Two learners equal in wealth and unequal in drift therefore price the same assay differently, and an allocation optimal for one is not optimal for the other.
Section 21.6 declined to say how a learner allocates its metered search, on the grounds that policy is one of the framework’s genuinely free parameters, and that remains right. But Proposition 26.3 narrows the space: an internalizing learner’s admissible policies are those that price an assay by both of the two parts above. That is not an attention mechanism. It is a constraint every attention mechanism for such a learner would have to satisfy, which is a weaker claim and one the framework can actually make.
It has set points. By construction, which is the point of the whole chapter.
One consequence falls out and should be recorded, because it is the first thing in this book that deserves the name.
Suppose \(\suff_i < 0\) and \(\suff_j < 0\) for two flavors, and suppose the enzyme graph offers no arbitrage-free conversion between them — the situation §24.10 calls incommensurability. Then no reallocation closes both shortfalls, and no exchange rate exists at which the learner could decide how much of one deficit is worth how much of the other. The learner is not confused and has not miscalculated. It is in a state which is, in the exact technical sense already available in Chapter 24, a conflict of values.
Where the conversion is arbitrage-free, the virtual token supplies the exchange rate and the conflict is merely a computation. So the framework distinguishes tractable trade-offs from genuine dilemmas, and the line between them is a property of the enzyme graph rather than of the learner’s psychology.
26.5 The price, and who pays it
None of this is free. An internalizing learner runs an instrument, holds a formula, and funds assays about itself out of the account those assays are about; Chapter 35 makes that charge precise. So the question that governs whether any of this happens is the one the chapter has been avoiding: when is looking worth it?
The answer is not available to the learner.
To decide whether to become internalizing, a learner would have to compare its prospects with and without a self-model. Both terms of that comparison are facts about a model it does not yet have. The deliberation would require the apparatus whose acquisition is being deliberated, and there is no order in which the learner could carry it out.
So internalization is not adopted. It is inherited or it is not, and what settles the question is which lineages there are.
26.5.1 Mortality first: what Buss showed
The argument we need has a precedent, and the precedent is worth rehearsing because the shape is the same and the biology is better known.
Von Neumann, asking what a self-reproducing machine must contain [16], isolated a requirement that looked paradoxical: the machine’s description has to be used in two incompatible ways. It must be copied uninterpreted, handed on as inert instructions, and it must be interpreted, read as a recipe for building the offspring’s working parts. One object, two uses, and the machine must keep them apart.
Definition 21.12 is that distinction, on a single object. The genome is \(@P\) and the phenotype is \(\ast @P\). The quotation does not reduce, does not age, and burns no tokens while held; the dereference is what is metered. Remark 21.17 adds the part that grammar cannot supply: what makes a reinstatement soma rather than offspring is provisioning, not syntax. We take no position on whether this constitutes a Weismann barrier; Remark 21.16 already declined that, and nothing here needs it.
Buss’s problem [15] is then immediate. A multicellular organism is a coalition of cells, every one descended from ancestors that did nothing but divide as fast as they could. What stops a fast-dividing line inside the body from doing what its ancestors did — outcompeting its neighbors, taking the whole, and destroying the individual in the process? Why is a body stable against its own quickest part?
In an unmetered theory, a replicating subterm whose cycle is strictly faster than its neighbors’ invades any parallel composition without bound. In a cost-accounted theory over a bounded inflow, a subterm drawing more per cycle than it earns runs for a number of cycles bounded by its endowment and then deadlocks by starvation. Its cumulative displacement of the whole is therefore bounded by that endowment.
The first clause is the absence of any rule that could stop it: in the free theory every enabled step is affordable and replication is unconditioned. The second is Definition 30.6 together with the conservation of the flavor: draw exceeding earnings is a strictly decreasing account, and an account below the charge of the next rewrite enables nothing.
That is Buss’s resolution with a mechanism under it. A configuration that sequesters the minting power into a germ line and lets its somatic parts be mortal admits no heritable takeover by a high-draw defector, while no configuration of the free theory resists one. Mortality is not tolerated in spite of bounding individual lives. It is selected because a bounded life is the precondition for there being a stable higher-level individual at all — a thing that can itself become a durable target of selection. Death is what makes the body possible.
26.5.2 And then: what internalization costs a lineage
Now run the same argument one level up.
The instrument, the held formula, and the per-cycle assay of Proposition 35.1 are somatic overhead. They are funded from the account that also funds foraging, maintenance and the endowment of offspring, and by Chapter 24 that account is bounded by what the medium supplies. So an internalizing lineage reproduces more slowly than an otherwise identical merely-surviving one, other things equal. In a placid world, other things are equal, and it loses.
What buys the overhead back is variance.
Let \(E\) be an ecology with per-cycle inflow of mean \(\mu\) and variance \(v\) at a locus, and let \(\kappa_{\mathrm{reg}}\) be the regulation cost of Proposition 35.1. Then there is a threshold \(v^\ast(\kappa_{\mathrm{reg}}, \mu)\), increasing in \(\kappa_{\mathrm{reg}}\), such that internalizing lineages are invaded by merely-surviving ones when \(v < v^\ast\) and are not when \(v > v^\ast\).
The mechanism intended is anticipation. A merely-surviving learner corrects when it fails to afford a rewrite, which is to say after the level has fallen. An internalizing learner reads a drift rate and can correct before, committing tokens to foraging while it still has tokens to commit. Where inflow is steady the two policies coincide and the second is simply dearer. Where inflow is ragged, the first spends part of its life below the charge of its own next step, and that is the interval in which the second is not merely more comfortable but alive.
We state this as a conjecture and not a theorem, and the honest reason is that the framework does not yet supply a stochastic inflow process to run it against. Chapter 13 supplies the machinery — propensities, a Gillespie unraveling, a well-posed simulator — and Conjecture 26.1 is stated in the form it would need to be checked in. That is an experiment this book has not run.
The shape of the argument is the reason it is worth making. The learner never evaluates the trade. It cannot, by Remark 26.8, and it does not need to. What evaluates the trade is the difference between lineages, run for long enough, in an environment that is or is not ragged.
That is also the answer to the question a reader may have been holding since Definition 26.2: where does \(\setpt\) come from? Not from the learner. A setpoint is a heritable parameter, carried in \(@P\) and therefore subject to variation and to Proposition 21.16’s recombination, and a lineage whose setpoint is badly placed is a lineage that either starves at the top of its comfort region or wastes at the bottom of it. Setpoints are selected in exactly the way body sizes are.
26.6 What is claimed, and what is not
Claimed: that an ecology of the kind Part Part II builds already contains everything needed for a learner to hold a goal about its own reserves; that when it does, a signed vectorial quantity appears which behaves in several specific respects like affect and which no previous construction in this book could express; that this quantity is derived rather than posited, which the thing it replaces was not; and that its cost is real and is paid by lineages rather than by individuals.
Not claimed: that there is anything it is like to be such a learner. Definition 26.4 is a structural analogue and the word chosen for it is a word with a great deal of freight. We use it deliberately — “negative valence signal” would be a way of not saying what we mean — but the reader should hold us to the structure and not to the connotation. Chapter 63 lists the explanatory gap among the things this book does not close.
Also not claimed: that the setpoint is unique, that \(\varepsilon\) has any principled value, or that Conjecture 26.1’s threshold has been computed. It has not; it has been located.
It is worth ending on the smallest version of the gain, because the chapter has been long.
Before this chapter, a learner in trouble and a learner at rest were distinguished by an observer with access to the ledger, and by nothing else. The learner itself had no state that differed. After it, the two differ in a quantity the learner holds, reads, and acts on — and the quantity is not its wealth.
Whether that is worth calling suffering is a question about words. That the ecology admits the distinction at all, and that the distinction costs something and is therefore not universal, is a question about structure, and the answer to that one is yes.