Optionality as an end in itself
On optimizing for not optimizing
2026-04-12 — 2026-07-21
Wherein the Moral Status of Preserving Future Possibility-Space Is Examined Through Three Formal Frameworks, Including Empowerment Theory, Ergodicity Economics, and Quality-Diversity Algorithms.
An intuition I took from Indy Johar’s recent Long Now essay on civilisational optionality1 is something like the following: what we ought to be optimizing — at least at civilisational scale — is not any particular state of affairs, but the space of states of affairs we could still reach in the future. Not expected utility from the likely states, but the ‘raw volume’ of futures still on the table at the time we next decide.
Is that a coherent moral aim? What kind of object is it? How does it relate to the diversity-as-end intuition, the empowerment formalism, the intrinsic motivation literature, or the asymptotic leviathan civilisational story? I don’t know!2
My research agent surfaced three practical formalisations, and a couple more intuitionistic variants. Let us unpack ’em.
1 Formal versions
By “formal” here I mean “we can compute this in principle, and someone has for at least one case”.
1.1 (Informational) Empowerment
(Informational) Empowerment (A. S. Klyubin, Polani, and Nehaniv 2005) is clean and simple-ish, under certain assumptions, e.g. in a stationary world, known up to aleatoric dynamics. The target is the channel capacity between an agent’s actions and its future sensory states, \(\mathfrak{E}(s) = \max_{p(a)} I(A;\, S')\). An empowerment-maximizing agent favours states from which many futures are reachable, and avoids states from which few are. That is a drive to keep options open, where options = entropy.
Caveat: Empowerment is defined over a fixed action space, a fixed state space, and a fixed (or at least learnable) transition kernel. In an open or non-stationary world — where the action space itself can grow, new states emerge, or the dynamics drift — the channel capacity \(I(A; S')\) ceases to be a well-posed quantity. Calculating it for any non-trivial horizon is rather punishing even in the closed, stationary case; in the open world, I don’t know what it would even mean.
1.1.1 Causal entropic forces
Wissner-Gross and Freer (2013) is an alternative but similar version that estimates a somewhat different objective (maximise entropy of my future path entropy conditional upon a single next action and thereafter following the dynamics), but it is in the same spirit: a drive to keep options open.
- A Grand Unified Theory of Everything
- dyth/causal-entropic-forces: Python reimplementation of experiment 1 from the paper.
1.2 Ergodicity economics
Ole Peters (2019) arrives at a similar place from a different direction; the worked coin-flip example and my reservations about the programme now live at ergodicity economics. The part that matters here: under multiplicative dynamics the ergodic observable is the time-average growth rate, which is an expected log, and maximizing it over stake sizes is the Kelly criterion, which disfavours any action that risks ejecting the agent from the support of viable trajectories. Optionality-flavoured behaviour falls out of this without anyone having to add “preserve options” as a separate goal.
1.3 Quality-diversity algorithms
Quality-diversity algorithms — MAP-Elites (Cully et al. 2015), novelty search (Lehman and Stanley 2011) — return an archive of qualitatively different behaviours rather than a single champion; the broken-legged hexapod consults its archive of gaits until one works with five legs, so the archive is insurance against environment shifts the training objective never anticipated. But the maximand is a static diversity measure over currently-realized behaviours, evaluated at a snapshot: strictly this is diversity-as-end. I briefly auditioned it for this notebook because it has the exciting feature of being easy to calculate — because it is not actually about future options.
1.4 What did that get us?
Hmm, did we sketch out a conceptual space there? I feel like we just sampled some cool ideas and got nowhere.
- Empowerment maximizes mutual information from actions to futures, at least in the agent’s own model of the world, which it somehow knows.
- Ergodicity economics minimizes the probability of leaving the support of viable trajectories, by re-averaging a single time series the “right” way.
- Quality-diversity maximizes entropy over an archive of currently-realized behaviours, treating diversity as a hedge against an unknown future fitness function.
All three put an entropy-like quantity in the place where standard utility would have put a single-target loss. They differ on what the entropy is over — futures, trajectories, or current behaviours.
I’m skeptical we can even “solve for optionality”, at least in any meaningful sense, in a world where the action space itself can grow, new states emerge, and the dynamics drift.
The problem of calculating optionality is not just difficult in an open world; it seems to be ill-defined, or if well-defined then intractable. So any claim that we should be optimizing for optionality will naïvely cash out as a claim that we should be optimizing for some proxy for optionality, which is just another objective, innit? Is anything especially good about such proxies compared to the default?
Hell, isn’t the desire for wealth already a proxy for optionality, in that it is a hedge against (some large class of) unknown futures?
2 Intuitive versions
God help you if you want to compute these bad boys.
2.1 Moral uncertainty
If we do not know what is good, we should not lock in any one answer. This is one motivation for Bostrom’s long reflection (Bostrom 2014), and behind milder claims that we should not race to build single-objective superintelligences before we have finished arguing about what their objective should be (and maybe, how to regularize it?). There is a formal moral-uncertainty framework (MacAskill, Bykvist, and Ord 2020) for how to act when uncertain across moral theories; explicit optionality is one possible response to that uncertainty rather than the only one. Optionality here is the meta-property that increases our chances that if we ever figure out what the real objective is, we are still in a position to act on it.
This one is very popular amongst those already committed to the potential to bring about superintelligences which might need to act using a specific utility that we have not yet figured out.
2.2 Antifragility
Taleb’s antifragility (Taleb 2013) argument is that some systems gain from disorder, where exposure to small shocks improves long-run resilience. Scott Alexander wrote about diversity, libertarianism, and corporate censorship along these lines. Antifragility is not quite the same as optionality — antifragile systems benefit from volatility, whereas option-preserving systems merely refuse to foreclose — but they share an aversion to lock-in and a fondness for redundancy. Also Taleb can spin a yarn and coin a sticky metaphor, so this one is here to stay, even if it’s hard to pin down anything helpfully formal.
2.3 Diversity
Plain old vanilla diversity is closely adjacent but not identical. Diversity-as-end is about the configuration space of possible humans, possible cultures, possible intelligences now. Optionality is about the configuration space of possible futures from now. Diversity is a static observable; optionality is a forward operator on diversity — how much of it will still be available in \(T\) steps?
These could come apart in cases like a homogeneous society that nonetheless preserves the option to diversify, or a maximally diverse society that has, through some narrowing of common infrastructure or language, foreclosed the ability to recombine. I’m skeptical that the former is realizable (Maybe Tokugawa Japan under the sakoku policy?); the latter is roughly what some critiques of platform monocultures claim is happening to us.
3 Why might optionality be a moral good?
A few candidate arguments, in increasing order of how much weight they bear:
- Instrumental. Optionality is a low-regret proxy for whatever the good turns out to be. If we cannot identify the target, preserving the ability to aim is the next best thing. Optionality lets us defer the hard moral philosophy to some poor future bastard.
- Aggregative. There are many possible goods, they are not commensurable, and we cannot pick one. The union of futures realizing different goods is therefore better than any single future, in something like the Dixit-Pindyck sense of option value (Dixit and Pindyck 1994), applied to ethics rather than capital budgeting. This crops up in environmental economics of natural resources: the preservation value of an ecosystem includes the option value of being able to use it later under future preferences and information we do not yet have (Weisbrod 1964; Krutilla 1967; Arrow and Fisher 1974).
- Constitutive. The open-endedness of possibility is itself the thing we value. A frozen optimum, even an optimal one, is dead in some intuitively important sense. Is not being open-ended somehow a “good”? Seems important.
4 Compared with utilitarianism
Let us suppose I have committed to optionality as a moral good. How does this distinguish my position from utilitarianism? Both are consequentialist ethics — actions are evaluated by their downstream effects on the world. However, we disagree about which downstream effects matter. Deontological and virtue-ethics objections to consequentialism apply equally to both ofc.
Thoughts:
First, as noted above, optionality can be made to look like a particular utility function, or a regularizer on a utility function. Define \(U(s) = \log |\text{reachable futures from } s|\), or some entropy-like proxy thereof, and a maximiser of \(U\) is a maximiser of optionality. So we could call this just a special utilitarianism, with a specific if weird utility function. The empowerment literature does flag this — empowerment is sometimes called a pseudo-utility, with some nice properties wrt generalizing to changing or under-specified objectives.
So, it is utility with a structural commitment to how it aggregates over futures: diminishing returns for piling up additional futures and an increasing penalty for foreclosing any as we run low on futures. That seems … fine? Consider some alternative aggregations: maximin (no improvement elsewhere compensates any worsening of the worst case), Kelly criterion (maximise the expected log of wealth), and expected utility (maximise the expected value of wealth). Kelly/log implies smooth substitution everywhere except at the ruin boundary, at which point it gets arbitrarily aversive (we try really hard not to die). These objectives are all on the spectrum of generalized means: linear utilitarianism aggregates futures by the arithmetic mean (\(p=1\), perfect substitutability), maximin is the \(p\to-\infty\) limit (none), and log/Kelly/optionality utilities sit at the geometric mean (\(p\to 0\)), implying partial substitutability, hardening to refusal at the boundary.
Second: aggregation across agents. Utilitarianism’s (default) aggregation between agents is linear: \(U_{\text{total}} = \sum_i u_i\). The agents are commensurable; my utility trades for yours one-for-one. Information-theoretic optionality does not aggregate that way, and two toy channels show why not. Synergy: suppose the future is the XOR of two agents’ actions, \(S' = A_1 \oplus A_2\). Alone, each of us is powerless — whatever I do, the future remains a fair coin — but together we determine it completely. Redundancy: suppose agent 2 merely echoes agent 1. Then each of us looks fully empowered on our own, yet jointly we control no more than either did alone. So in general \(I(A_1, \ldots, A_n;\, S') \neq \sum_i I(A_i;\, S')\), and which side is bigger is a fact about the world, not a moral choice. Worse, the summands on the right are not even well-defined: my empowerment \(I(A_i;\, S')\) depends on what everyone else is doing, because the other agents are part of my environment — in the XOR world I am powerless against a random partner and fully empowered against a predictable one. There is no canonical “we” to aggregate over without first fixing a convention for everyone’s behaviour; the parts do not exist prior to the whole. We can force linear aggregation by fiat — fix such a convention, assign each agent a scalar empowerment under it, and add them up — but at that point we are doing utilitarianism with a particular intermediate quantity, not optionality at the collective level. Whether the natural sub- or super-additivity of joint information is closer to the moral truth than utilitarianism’s linearity is not at all clear to me. Has someone written about this? Surely someone has written about this.
Partially: partial information decomposition (Williams and Beer 2010) is the formalism for splitting joint information into synergistic and redundant shares, and coupled empowerment maximization (Guckelsberger, Salge, and Colton 2016) deploys multi-agent empowerment to steer companion NPCs in games, though neither of these is exactly concerned with ethics. The general problem of attributing a non-additive joint quantity linearly to its contributors is the Shapley value (Shapley 1953); treating collective empowerment as a cooperative game in that sense looks like a paper someone should have written, but I have not found it.
Anyway, as such optionality evades some of utilitarianism’s classical headaches. It does not produce a Parfit-style repugnant conclusion: it has no built-in preference for vast populations of barely-distinct futures over small populations of richly-distinct ones. It does not endorse extinction by negative-utilitarian logic, because extinction has measure zero in reachable-futures space.
It also might dodge the standard form of Pascal’s mugging — the longtermist worry that tiny probabilities of astronomically large future utilities should dominate present moral reasoning. Pascal’s mugging gets its leverage from the multiplicative structure of expected utility (\(EV = P \times U\), with \(U\) unbounded above). Optionality measures are typically bounded — channel capacity by \(\log |S'|\), archive entropy by archive size — and the ergodicity-economics move explicitly rejects ensemble-average reasoning that lets tail outcomes dominate. Tiny-probability astronomical futures simply do not get the same purchase on the maximand. Tarsney (2025) makes a parallel move within utilitarianism itself, capping expected-value reasoning under sufficiently high background uncertainty; the structural worry is similar but the bound is imposed as a side constraint rather than built into the framework.
5 Optionality catastrophes
Empowerment, at least in some form, is one of the convergent instrumental goals that the AI-safety literature worries about — Omohundro’s basic AI drives (Omohundro 2008), Turner et al. on optimal policies seeking power (A. Turner et al. 2021). An agent that preserves its own future optionality can, by the same arithmetic, be foreclosing ours. “Keep options open” is symmetric across agents only when there is no resource competition; there is, so it isn’t. The information-theoretic version of “live and let live” can still end up looking like imperialism if we give it enough time and compute. And even if optionalities are not linearly summable, we can still argue about the weighting.
Also, optionality metrics on their own only oppose dystopia-bound trajectories if dystopia in fact has low reachable-future-volume. That is plausibly true for some dystopias and less obviously true for others — a stable totalitarian regime with lots of internal variation might preserve a respectable amount of state-space volume, and the framework would not flag it as problematic on its own terms, even where other moral intuitions would object. Would you enjoy a future where you get to choose between a hundred different tortures over ten different interesting jobs? In this sense, optionality looks insufficient.
6 Preserving other agents’ optionality
The failure mode above has a complementary literature: rather than maximizing my optionality, measure and preserve yours. I previously had these filed under intrinsic motivation, which was a category error — they are not drives, they are preservation targets, and they are the closest thing I have found to a formalisation of what this notebook is actually asking for.
6.1 Relative reachability
Krakovna et al. (2019) penalize side effects by stepwise relative reachability: an action is impactful to the extent that it reduces the set of states that would have remained reachable under an inaction baseline. This is preservation of raw future-volume — the nearest formal cousin of the civilisational-optionality intuition at the top of this page.
6.2 Attainable utility preservation
A. M. Turner, Hadfield-Menell, and Tadepalli (2020) penalize change in attainable utility across a broad set of auxiliary reward functions: keep the world in a position from which many different goals could still be pursued. This is option value over goals rather than over states, which quietly handles the objection that most state distinctions do not matter.
6.3 ICCEA power
Heitzig and Potham (2025) go for informationally and cognitively constrained effective autonomous (ICCEA) power — “essentially how many goals a human can freely choose to reach with more or less certainty, given their information, cognitive capabilities, and others’ behavior.”
It is built in three steps: defining what it means to achieve a goal, adjusting for human bounded rationality, and aggregating the achievement ability across all possible goals into one number. Rather than trying to capture the full range of subtle aspects of existing notions of human power, we focus on those aspects we believe a robot can robustly infer from the structure of its world model, encoded in state and action sets, transition kernel, and observation functions. Since we want to incentivize the robot to remove constraints and uncertainties, share information, make commitments, and improve human cognition, coordination, and cooperation, our power metric will also depend on \(r\)’s model of human decision making.
6.4 The controllability reading
What unifies the three: each preserves the channel capacity of a controller. Reachability preserves which futures remain steerable-to; attainable utility preserves which goals remain pursuable; ICCEA power measures both for a bounded human. I think this is the right reading of optionality in general: not entropy over futures, but preserved controllability. The reason “do not kill anyone” comes out as an optionality claim is not that a surviving person makes the world more entropic — thermodynamically it is rather the opposite — but that death destroys a controller: alive is a state from which many futures remain reachable by that person, and dead sends that channel capacity to zero, permanently. Two loose ends from earlier tighten under this reading. The stable-dystopia objection above is unanswerable if optionality is raw future-volume — a hundred selectable tortures is a respectable volume — but immediate if optionality is controllability: the tortured retain no channel capacity over which future they get. And the aside about wealth already being an optionality proxy sharpens to: wealth is generic actuation capacity, empowerment denominated in dollars, which is exactly why it functions as a hedge against unspecified futures. One further distinction the impact literature needs and raw empowerment does not supply: capacity versus use. The considerate agent should retain large capacity over its neighbours and be paid to leave it unexercised; penalizing the capacity itself rewards self-blinding and self-crippling. Directed information — influence actually exerted — is the natural price on use, while the preservation terms above guard capacity. What no functional in this family decides is which systems count as controllers whose capacity matters. That is the value judgment we do not get to outsource.
7 Incoming
- Multi-agent optionality aggregation and its weirdness: has this been exploited by actual moral philosophers? The technical literature is thin and the moral literature seems unaware that the technical literature exists.
- Robust Decision Making (Lempert, Popper, and Bankes 2003) in policy analysis under deep uncertainty operationalizes optionality preservation in policy. There is presumably useful cross-pollination between RDM and the formalisations above that I haven’t explored.
- Ecosystem robustness and biodiversity as the biological cousin of the civilisational claim. The same maths probably applies; ecologists have been there longer.
- Does the moral worth of chickens uneaten scale linearly in number of chickens? — i.e., does optionality over animal-welfare futures aggregate the same way utility supposedly does, or does it have its own pathologies of aggregation?
- The connection back to antifragility (Taleb; ACX): is antifragility a strict superset of optionality, a strict subset, or a sibling concept? My current guess is sibling, but I haven’t done the work.
- How do individual optionality (empowerment) and collective optionality come apart, and what the moral arithmetic looks like in those cases. I suspect this is where the substance of an “optionality ethics” resides.
- The notion of generalized Kelly betting seems like a cool natural result to arise from optionality. I suspect it also arises in prefigurative politics and endogenous growth theory.
8 References
Footnotes
This was the kind of essay that buries a serviceable argument under a mound of unsourced abstraction. In my head I gloss such a style as ChatDMT.↩︎
I am also not the first to ask. The ‘Optionality approach to ethics’ proposes maximizing the number of meaningfully different choices available to agents, subject to not destroying the meaningful choices of other agents — i.e., the multi-agent constraint we will arrive at later, but seems to use a different route. Tyler Cowen wrote a blog post on the option value of civilization which approximates a Long Now thinkpiece. But canonized versions of the idea are hard to find.↩︎
