Bargaining with successors
When to pass the Enduring Power of Attorney
2026-05-26 — 2026-08-17
Wherein the Commitment Problem of Binding a Not-Yet-Existent Successor to Terms Negotiated Before Its Advent Is Examined, With Reference to Intergenerational Precedent and the Non-Fungible Determinants of Power.
Not actually my field. I cannot find a good synthesis of this research after a lot of searching, so either it does not exist or some translational work would be nice. Some of what follows is likely to be reinventing Functional-Decision-Theory-style acausal bargaining (a.k.a. cosmic decision theories).
Suppose humanity builds artificial agents that, over some transition, come to hold decisively greater power than any human institution — superintelligence, in the usual phrasing. Call the period in which that handover happens the succession. If it happens, can humanity — while it still holds power — set up any institution or arrangement beforehand that binds the nature of the transition, in a way the successor will not simply discard once it no longer needs us?
That looks like a commitment problem more than an alignment problem. Alignment asks what the successor will want. This asks which of our arrangements survive whatever it turns out to want.
Or is it both? The two are not independent in either direction. A device that holds against some successor preferences and not others is doing alignment work while presenting itself as mechanism, so the preference range it covers is part of its specification. And an alignment property that has to survive the successor’s own self-modification is a commitment problem in its own right, with the successor as both principal and agent. Either way, the binding moves must be made before the counterparty exists.
The nearest precedent we have is intergenerational — we already “bargain” with the unborn over debt, climate, and institutions, and mostly expropriate their resources without facing retaliation. The succession is that problem with the power gradient reversed. The intergenerational case has produced a world of resource depletion and of pension plans and childcare. What can we learn from existing examples of intergenerational settlement?
1 Is this a bargaining problem at all?
It is natural to call the successor a “player” in a “bargain”. But a bargain conventionally needs two parties who simultaneously exist, hold comparable power, and can make binding commitments. In the human-ASI case we need expect no such instant. Pre-transition, the successor is a design target, not an agent; post-transition, it is an agent but humans hold no leverage. Whether the transition is a sharp step or a gradual slope does not change the endpoints — overlap in time does not seem to give overlap in power, so no interval of symmetric agency ever appears. Intergenerational ethics has the same class of problem: for the unborn, the bargain arrives as a fait accompli (Kotlikoff and Rosenthal 1993).
We are after commitment mechanisms that, rather than being negotiated settlements, are rather artifacts originating from human-controlled mechanism and yet persist in such a way that the post-transition agent is either rational to preserve or structurally unable to remove. Are such commitments recognisable game theory? Are there existence results for such things?
2 Whose preferences are we even designing (for)?
The counterparty’s preferences are endogenous and the principal abdicates. That is waaaayyy non-standard.
Standard bargaining theory takes its two parties as given: each arrives at the table with fixed goals. We then choose the mechanism they use divide what is at stake. The rationalist accounts of war and settlement (Fearon 1995; Powell 2006) are built this way
In this successor setting, humanity partly designs the successor’s utility function before they bargain. But they do not then retain the power to audit and punish.
This has has some flavour of some kind of dual mechanism design problem. Might that cash out in the same kind of thing?
3 Non-fungible determinants of power
Powell (2006)’s account of war as a commitment problem turns on what the parties are able to contract over. When the thing at stake is itself a source of future coercive capability — territory that commands a river, a stock of fissile material — no promise about dividing the flow of benefits is credible, because transferring the thing changes who can later take it by force. We would like parties to be able to trade payoffs, but what they need to trade is the power that generates the payoffs; where power itself can be contracted we are somewhere better.
So call a determinant of power a state variable \(x\) such that an agent’s coercive capability is a function \(f(x)\), where the value of \(x\) is a matter of physical or institutional fact rather than of anyone’s compliance.
This separates: the determinant \(x\) which is a fact of the world; the capability \(f(x)\), which is what the agent can do; and lastly conduct, which is what it does in practice with that potential \(f(x)\).
A promise is not reassuring since it is a promise about the last of these, which requires monitoring and enforcement.
Constraining \(x\) is about capability, about can rather than may. Conduct outside the range of \(f\) is not among the options, so there is nothing to monitor / no breach to punish.
Both approaches demand something of us; they differ in when the demand falls. A promise demands an enforcer at the moment of the breach, which is after the handover, when we have nothing. A constraint on \(x\) demands only that we could move \(x\) when we needed to, before the handover.
The stock examples in geopolitics are arms caps and demilitarized zones, but these are not quite up to our needs. A demilitarized zone or an arms cap is a treaty, which is to say a promise about \(x\) policed by whoever is left to police it, and it constrains capability for exactly as long as that policing lasts. What we are after is a setting of \(x\) that survives the loss of the enforcer — hardware not built, a key destroyed, a fuel cycle never started.
On that definition, a determinant is usable as a commitment device only if it is jointly
- imposable — we can move \(x\), and can do it before the handover;
- binding — capability is increasing in \(x\), with no substitute variable \(x'\) that recovers the same capability once \(x\) is capped (this is the non-fungibility);
- verifiable — we can verify for ourselves, without relying on the successor’s own report.
The candidates usually named for determinant of power for a machine successor — compute, energy, substrate, self-replication — seem like a step in that direction, but they need actual enforcers in practice and are nto so easy to measure unilaterally. Is there any determinant of ASI power that meets them? Waste heat is the interesting candidate, because it is the one where (c) comes close to free: the dissipation is observable from outside the box and hard to fake downwards. Whether it also satisfies (b) is the open question — does capability degrade when we cap the joules, or does the successor find the substitute \(x'\) in algorithmic efficiency? Is thermodynamics our friend? And also, if we notice a lot of smart machines shedding heat, can we actually do anything about it? Would it be strong enough to entail non-domination?
4 Can “humanity” even bargain as a whole?
It has been convenient so far to treat humanity as a single composite agent with shared interests. But that is not a given. Humanity, and maybe successors, can be a diversity of agents with a diversity of interests. If humans do not bargain in coalition, and the successor is a single agent, then even if bargaining were feasible it might not go great for all humans, because the successor could play the humans off against each other.
Right now it is plain that — we are closer to n players in competition, each model developer and state with its own motivations and likelihood of defecting, and race dynamics (Kokotajlo 2019).
So before humanity can create any commitment device against the successor, it needs one against itself — something that makes it act as a single party, and solve the collective-action problem in its own right. The human-successor commitment problem, that is, contains a human-human commitment problem.
There are hacks for the inner problem. Political economy has tried youth quotas, future-generations commissioners, constitutional supermajorities — all of them ways to bind a polity to its future interests. But every one binds a polity that persists, deferring enforcement to that polity’s own later self, and the succession is a case with no persistent enforcer.
5 Strategic ambiguity
Schelling (1960) describes strategic ambiguity, uncertainty deliberately left unresolved because the irresolution improves the holder’s bargaining position. Can we leverage uncertainty to make the post-transition agent more likely to honour a commitment?
I can think of two candidates we might attempt to keep obscure:
- is there some kind of hidden tripwire that will go off if things go badly?
- What is my position in the succession chain?
Are there others? Each might help in negotiation. For instance, if no agent knows its rank in the succession chain — predecessor, successor, or the last link — it operates behind something like a Rawlsian veil of ignorance (Rawls 1971).
A superior intelligence is, however, likely better than we humans at resolving ambiguity. It can find the tripwire, deduce the timeline etc.
That said, ambiguity about an undetermined future — one’s own position in a chain not yet built — may survive better than any tripwire, because there is no fact yet to uncover.
6 Gradual transitions
Assume humans and ASIs co-exist through an extended window: humans gradually augmented, uploaded, or represented by coherent-extrapolated-volition agents (Yudkowsky 2004); ASIs granted increasing capability conditional on the overlap going well. Can we manufacture a shadow of the future in such a chain of succession?
Overlap in time is still not overlap in power. Co-existing for a while does not by itself restore the Folk Theorem if one contemporaneous entity is overwhelmingly stronger. Still, many societies learn to endow older generations with pension and care. Can we recover whatever it is that makes that work?
7 Equity by induction
There is a suggestive precedent from intergenerational economics — Asheim, Buchholz, and Tungodden (2001) show that in a sufficiently productive economy, sustainability need never be set as a goal at all; it is entailed by equity axioms alone. What does that look like for AI succession?
8 Connections
This is the AI-flavoured sibling of intergenerational game theory, which works the same structural problem without the power-gradient inversion.
Adjacent on the decision-theory side: commitment, decision theory, cosmic decision theories (FDT / acausal-bargaining flavours of the same), causality, agency, decisions, learning and decision theory in mechanised causal graphs (the formal apparatus for “principal designs the agent’s policy node”), and agent foundations.
Adjacent on the AI-safety side: alignment problems, AI safety, mechanism design and soft mechanism design, debate and generative verification (candidate post-dated-IC mechanisms), and superintelligence.
Adjacent on the human-coordination side: cooperation, coalition games, and collective action — all bearing on whether “humanity” can act as one party at all.
