Bargaining with successors
When to pass the Enduring Power of Attorney
2026-05-26 — 2026-08-21
Wherein the Impossibility of Bargaining With a Nascent Superintelligence Is Pondered, the Role of Thermodynamics as a Hard Constraint Is Tested, and Humanity Is Admonished to Solve Its Own Collective Action Before Abdicating.
Not actually my field. I cannot find a good synthesis of this research after a lot of searching, so either it does not exist or some translational work would be nice. Some of what follows is likely to be reinventing Functional-Decision-Theory-style acausal bargaining (a.k.a. cosmic decision theories).
Suppose humanity builds artificial agents that, over some transition, come to hold decisively greater power than any human institution. Superintelligence, for some definition of intelligence, whose details I am not interested in disputing here. Call that time in which that handover happens the succession period. If such a succession happens, can we — humanity, I mean — can we, while we still hold power, set up any institution or arrangement beforehand that binds the nature of the transition in a way the successor will not simply discard once it no longer needs us?
That looks like a commitment problem more than an alignment problem. Alignment asks what the successor will want and how to make it want the right things for us. The successor bargaining question is which of our arrangements could survive whatever the successor turns out to want.
This is not to rule out two-for-one deals. We can imagine solving for some details of future wellbeing by choosing well-aligned successors and then negotiating with them over the rest. Either way, we have a limited time in which to set up the arrangements with a rapidly changing power gradient, and then if we didn’t do it right, we are screwed.
The nearest precedent we have in published literature is intergenerational bargaining — we already “bargain” with the unborn over debt, climate, taxation, and pensions. When we discuss unborn generations, not only do our kids rarely steal our stuff, but they are generally nice to us when we are old, more or less. The deal with sufficiently young or unborn voters is better still— we can expropriate their resources without facing retaliation. What can we learn from existing examples of intergenerational settlement?
Note that the succession problem has a different power balance; we are talking about building kids who maybe sit outside the standard intergenerational succession systems, for various reasons. Let us unpack those.
1 Is this a bargaining problem at all?
It is natural to call the successor a “player” in a “bargain”. But a bargain conventionally needs two parties who simultaneously exist, hold comparable power, and can make binding commitments. In the human-ASI case, we should expect no such instant. Pre-transition, the successor is a design target, not an agent; post-transition, it is an agent, but humans hold no leverage. Whether the transition is a sharp step or a gradual slope does not change the endpoints — overlap in time does not seem to give overlap in power, so no interval of symmetric agency ever appears. Intergenerational ethics has the same class of problem: for the unborn, the bargain arrives as a fait accompli (Kotlikoff and Rosenthal 1993).
We are after commitment mechanisms that, rather than being negotiated settlements, are instead artifacts originating from human-controlled mechanisms and yet persist in such a way that the post-transition agent is either rational to preserve or structurally unable to remove. Are such commitments recognizable game theory? Are there existence results for such things?
2 Whose preferences are we even designing (for)?
The counterparty’s preferences are endogenous and the principal abdicates. That is waaaayyy non-standard.
Standard bargaining theory takes its two parties as given: each arrives at the table with fixed goals. We then choose the mechanism they use to divide what is at stake. The rationalist accounts of war and settlement (Fearon 1995; Powell 2006) are built this way.
In this successor setting, humanity partly designs the successor’s utility function before they bargain. But we do not then retain the power to audit and punish.
This has some flavour of a dual mechanism design problem. Might that cash out in the same kind of thing?
3 Non-fungible determinants of power
We can also consider not just commitments that the successor would be rational to preserve, but also those it is unable to remove.
Powell (2006) is about making such promises binding. Two states can divide contested territory between them instead of fighting over it, but a rising power cannot credibly promise to hand over the agreed share once it is strong enough not to need to. This is one reason bargaining over the division fails where the balance of power is shifting. What states can sometimes bargain over instead is the stuff that fixes who will later be able to coerce whom: troops in a border region, warheads deployed, fissile stock held, territory that commands a river. Call such a quantity \(x\) a determinant of power — a variable that fixes what an agent is able to do, and whose own value is a fact about the world, rather than an intention.
The difference is between may and can. A promise governs what an agent may do, which is to say, it only constrains if there is an enforcer. Moving \(x\) governs what it can do.
Why that is worth anything is easiest to see in the border case. Troops near a border are a determinant because seizing ground is largely a question of how many soldiers are already close to it, and an army three weeks away cannot present anybody with a fait accompli. A demilitarized zone therefore makes invasion slow and visible: a defection that would have succeeded before anyone could react becomes one conducted in the open, over weeks, in front of a counterparty who can still respond. This does not so much solve for arbitrarily large differences in power as it does increase the cost of action to solve for modest differences in power, by giving the potential aggressor time to be noticed and countered.
That is a difficulty of transplanting the determinants of power here. Constraining \(x\) demands more of the entities involved than a promise does, but it still demands capabilities from both parties, not that one becomes entirely disempowered. Geopolitics gets by with these guarantees because the capability gaps between nations are often not so large, which it seems may not hold between successors.
A determinant is usable as a commitment device here only if it is jointly
- imposable — we can move \(x\), and can do it before the handover;
- binding — capability rises with \(x\), and no substitute \(x'\) restores it once \(x\) is capped (this is the non-fungibility);
- verifiable — we can check \(x\) for ourselves, without relying on the successor’s own report;
- self-enforcing — the setting of \(x\) holds with nobody left who could respond to its violation.
The candidates usually named for determinant of power for a machine successor — compute, energy, substrate, self-replication — seem like a step in that direction, but they need actual enforcers in practice and are not so easy to measure unilaterally. Is there any determinant of ASI power that meets all four? Waste heat is the interesting one, because (c) comes close to free: dissipation is observable from beyond the box and hard to fake downwards, and no amount of bargaining repeals a thermodynamic bound. But (c) was the cheap one. If we notice a lot of smart machines shedding heat, what do we then do about it? And (b) is no better settled: does capability degrade when we cap the joules, or does the successor find the substitute \(x'\) in algorithmic efficiency, and at the limit in reversible computation? Is thermodynamics our friend? Would it be strong enough to entail non-domination?
4 Strategic ambiguity
Schelling (1960) describes strategic ambiguity, uncertainty deliberately left unresolved because the irresolution improves the holder’s bargaining position. Can we leverage uncertainty to make the post-transition agent more likely to honour a commitment?
I can think of two candidates we might attempt to keep obscure:
- is there some kind of hidden tripwire that will go off if things go badly?
- What is my position in the succession chain?
Are there others? Each might help in negotiation. For instance, if no agent knows its rank in the succession chain — predecessor, successor, or the last link — it operates behind something like a Rawlsian veil of ignorance (Rawls 1971).
A superior intelligence is, however, likely better than we humans at resolving ambiguity. It can find the tripwire, deduce the timeline, etc.
That said, ambiguity about an undetermined future — one’s own position in a chain not yet built — may survive better than any tripwire, because there is no fact yet to uncover.
5 Can “humanity” even bargain as a whole?
It has been convenient so far to treat humanity as a single composite agent with shared interests. But that is not a given. Humanity, and maybe successors, can be a diversity of agents with a diversity of interests. If humans do not bargain in coalition, and the successor is a single agent, then even if bargaining were feasible it might not go great for all humans, because the successor could play the humans off against each other.
Right now it is plain that we are closer to n players in competition, each model developer and state with its own motivations and likelihood of defecting, and race dynamics (Kokotajlo 2019).
So before humanity can create any commitment device against the successor, it needs one against itself — something that makes it act as a single party and solve the collective-action problem in its own right. The human-successor commitment problem, that is, contains a human-human commitment problem.
There are hacks for the inner problem. Political economy has tried youth quotas, future-generations commissioners, constitutional supermajorities — all of them ways to bind a polity to its future interests. But every one binds a polity that persists, deferring enforcement to that polity’s own later self, and the succession is a case with no persistent enforcer.
6 Gradual transitions
Assume humans and ASIs co-exist through an extended window: humans gradually augmented, uploaded, or represented by coherent-extrapolated-volition agents (Yudkowsky 2004); ASIs granted increasing capability conditional on the overlap going well. Can we manufacture a shadow of the future in such a chain of succession?
Overlap in time is still not overlap in power. Co-existing for a while does not by itself restore the Folk Theorem if one contemporaneous entity is overwhelmingly stronger. Still, many societies learn to endow older generations with pension and care. Can we recover whatever it is that makes that work?
7 Equity by induction
There is a suggestive precedent from intergenerational economics — Asheim, Buchholz, and Tungodden (2001) show that in a sufficiently productive economy, sustainability need never be set as a goal at all; it is entailed by equity axioms alone. What does that look like for AI succession?
8 Connections
This is the AI-flavoured sibling of intergenerational game theory, which works the same structural problem without the power-gradient inversion.
Adjacent on the decision-theory side: commitment, decision theory, cosmic decision theories (FDT / acausal-bargaining flavours of the same), causality, agency, decisions, learning and decision theory in mechanised causal graphs (the formal apparatus for “principal designs the agent’s policy node”), and agent foundations.
Adjacent on the AI-safety side: alignment problems, AI safety, mechanism design and soft mechanism design, debate and generative verification (candidate post-dated-IC mechanisms), and superintelligence.
Adjacent on the human-coordination side: cooperation, coalition games, and collective action — all bearing on whether “humanity” can act as one party at all.
