Agency wat

Dan MacKinlay

2026-07-13

A talk for Human-aligned AI Summer School 2026.

Terminology note

  • “agent” is a terminology tarpit.

  • other tarpits we leave out of scope for now

    • alignment
    • empowerment
    • flourishing
    • rationality

Aligning what to what?

\[ \underbrace{\operatorname{align}}_{\text{not covered}}( \mathbf{A}, \mathbf{B} ) \]

Reasons to formalize agency

  1. We do it anyway, implicitly, when we talk about aligning
  2. In practice, modernism solves for something, so maybe agency is better than the alternatives (?)
  3. It seems to be important in our notion of moral patienthood

tl;dr agency ends up playing implicit normative and descriptive roles regardless, but it is not clear exactly what it is

Agents seem to be the entities in alignment problems

  • \(\operatorname{align}(🧑, 👶)\)
  • \(\operatorname{align}(🧑, 🧑)\)
  • \(\operatorname{align}(🧑, 🤖)\)
  • \(\operatorname{align}(🤖, 🤖)\)
  • \(\operatorname{align}(🧑, 🏟️)\)
  • \(\operatorname{align}(🪨, 🪨)\)

Implicit theories of agency

Many modes of AI risk have an implicit theory of agency, e.g.

Implicit theories of human agency

Pre-AI fields care a lot about agency

Models of agency to be confused about

  • decision theory — VNM / EV maximization · game theory
  • machine learning — RL / POMDP · AIXI · intelligence-is-optimization · agents-vs-devices · Discovering Agents · LLM + harness · simulators / role-play
  • information theory & cybernetics — cybernetic control · empowerment (Klyubin–Polani) · info-theoretic individuality
  • physics & biology — dissipative / autopoietic · active inference · enactivist / biogenic
  • psychology — empirical human psychology
  • law & polity — forensic (fit to be punished) · group / institutional
  • folk — “high agency” (Bay Area)
  • … a fog of others

Agency as a cluster concept

For a competent, adult human, these all coincide.

For machines, not so much.

Nor for superorganisms.

Axes of variation

  • boundary — agent/environment cut: primitive ↔︎ derived
  • multi-agent — solipsistic ↔︎ multi-agent-first
  • observer-relativity — agency intrinsic ↔︎ ascribed
  • boundedness — unbounded ideal ↔︎ substrate-bound
  • warrant — descriptive (is) ↔︎ normative (ought)

Is⇄ought crossings

We implicitly move between normative and descriptive all the time.

Agency as target

  • It seems like we generally think of agency as a good
  • High modernism works best if you can measure a good
  • Can agency be made objective enough to support this role?
  • Maybe agency is even a natural kind?

Fragility of value, agency edition

Specify “human agency” slightly wrong → agency-theatre

Existing challenges

Existing challenges

  • coercive control
  • addiction
  • attention capture

These all push the boundaries of already-existing intuitive notions of agency.

Existing challenges

  • What even is life?
  • In what senses are collectives agents?
  • Better metrics to track some good we care about

Future applications

  • Gradual disempowerment
  • Moral worth of machine intelligences

References

Arrow. 1951. Social Choice and Individual Values.
Arrow, and Debreu. 1954. Existence of an Equilibrium for a Competitive Economy.” Econometrica.
Bandura. 1977. Self-Efficacy: Toward a Unifying Theory of Behavioral Change.” Psychological Review.
Barandiaran, Di Paolo, and Rohde. 2009. Defining Agency: Individuality, Normativity, Asymmetry, and Spatio-Temporality in Action.” Adaptive Behavior.
Beauchamp, and Childress. 2019. Principles of Biomedical Ethics.
Becker, and Murphy. 1988. A Theory of Rational Addiction.” The Journal of Political Economy.
Boyd. 1991. Realism, Anti-Foundationalism and the Enthusiasm for Natural Kinds.” Philosophical Studies.
Bruineberg, Dołęga, Dewhurst, et al. 2022. The Emperor’s New Markov Blankets.” Behavioral and Brain Sciences.
Camara. 2021. “Computationally Tractable Choice.”
Carroll, Foote, Siththaranjan, et al. 2024. AI Alignment with Changing and Influenceable Reward Functions.”
Coase. 1937. The Nature of the Firm.” Economica.
Debreu. 1974. Excess Demand Functions.” Journal of Mathematical Economics.
Dennett, Daniel C. 1991. Real Patterns.” The Journal of Philosophy.
Dennett, D. C. 1998. The Intentional Stance. A Bradford Book.
Du, Tiomkin, Kiciman, et al. 2020. AvE: Assistance via Empowerment.” In Advances in Neural Information Processing Systems.
England. 2013. Statistical Physics of Self-Replication.” The Journal of Chemical Physics.
Flint. 2020. The Ground of Optimization.”
Frankfurt. 1971. Freedom of the Will and the Concept of a Person.” The Journal of Philosophy.
Gruber, and Kőszegi. 2001. Is Addiction ``Rational’’? Theory and Evidence.” The Quarterly Journal of Economics.
Gustafsson. 2022. Money-Pump Arguments. Elements in Decision Theory and Philosophy.
Hadfield-Menell, and Hadfield. 2018. Incomplete Contracting and AI Alignment.”
Harsanyi. 1955. Cardinal Welfare, Individualistic Ethics, and Interpersonal Comparisons of Utility.” Journal of Political Economy.
Hart. 1968. Punishment and Responsibility: Essays in the Philosophy of Law.
Hayek. 1945. The Use of Knowledge in Society.” The American Economic Review.
Hedden. 2015. Time-Slice Rationality.” Mind.
Hubinger, Merwijk, Mikulik, et al. 2021. Risks from Learned Optimization in Advanced Machine Learning Systems.”
Hutter. 2005a. The Universal Algorithmic Agent AIXI.” In.
———. 2005b. Universal Artificial Intelligence: Sequential Decisions Based on Algorithmic Probability. Texts in Theoretical Computer Science.
Hutter, Quarel, and Catt. 2024. An Introduction to Universal Artificial Intelligence.
janus. 2022. Simulators.”
Jensen, and Meckling. 1976. Theory of the Firm: Managerial Behavior, Agency Costs and Ownership Structure.” Journal of Financial Economics.
Kasirzadeh, and Gabriel. 2023. In Conversation with Artificial Intelligence: Aligning Language Models with Human Values.” Philosophy & Technology.
Kenton, Kumar, Farquhar, et al. 2023. Discovering Agents.” Artificial Intelligence.
Klyubin, Alexander S., Polani, and Nehaniv. 2005. All Else Being Equal Be Empowered.” In Advances in Artificial Life. Lecture Notes in Computer Science.
Klyubin, A.S., Polani, and Nehaniv. 2005. Empowerment: A Universal Agent-Centric Measure of Control.” In 2005 IEEE Congress on Evolutionary Computation.
Korsgaard. 2009. Self-Constitution: Agency, Identity, and Integrity.
Krakauer, Bertschinger, Olbrich, et al. 2020. The Information Theory of Individuality.” Theory in Biosciences.
Kulveit, Douglas, Ammann, et al. 2025. Gradual Disempowerment: Systemic Existential Risks from Incremental AI Development.”
Legg, and Hutter. 2007. Universal Intelligence: A Definition of Machine Intelligence.” Minds and Machines.
Levin. 2022. Technological Approach to Mind Everywhere: An Experimentally-Grounded Framework for Understanding Diverse Bodies and Minds.” Frontiers in Systems Neuroscience.
List. 2019. Levels: Descriptive, Explanatory, and Ontological.” Noûs.
List, and Pettit. 2011. Group Agency: The Possibility, Design, and Status of Corporate Agents.
Locke. 1689. An Essay Concerning Human Understanding.
Maturana, and Varela. 1980. Autopoiesis and Cognition: The Realization of the Living. Boston Studies in the Philosophy of Science.
Mohamed, and Rezende. 2015. Variational Information Maximisation for Intrinsically Motivated Reinforcement Learning.” In Proceedings of the 28th International Conference on Neural Information Processing Systems - Volume 2. NIPS’15.
Ngo. 2019. Coherent Behaviour in the Real World Is an Incoherent Concept.”
Nussbaum. 2011. Creating Capabilities: The Human Development Approach.
Omohundro. 2008. The Basic AI Drives.” In Proceedings of the 2008 Conference on Artificial General Intelligence 2008: Proceedings of the First AGI Conference.
Orseau, McGill, and Legg. 2018. Agents and Devices: A Relative Definition of Agency.”
Petersen. 2023. Invulnerable Incomplete Preferences: A Formal Statement.”
Prigogine, and Stengers. 1984. Order Out of Chaos: Man’s New Dialogue with Nature.
Ray. 2007. A Game-Theoretic Perspective on Coalition Formation. The Lipsey Lectures.
Ray, and Vohra. 2015. Coalition Formation.” In Handbook of Game Theory with Economic Applications.
Riker. 1962. The Theory of Political Coalitions.
Rogeberg. 2004. Taking Absurd Theories Seriously: Economics and the Case of Rational Addiction Theories.” Philosophy of Science.
Rosas, Geiger, Luppi, et al. 2024. Software in the Natural World: A Computational Approach to Hierarchical Emergence.”
Rotter. 1966. Generalized Expectancies for Internal Versus External Control of Reinforcement.” Psychological Monographs: General and Applied.
Ryle. 1949. The Concept of Mind.
Salge, and Polani. 2017. Empowerment as Replacement for the Three Laws of Robotics.” Frontiers in Robotics and AI.
Sen, Amartya K. 1977. Rational Fools: A Critique of the Behavioral Foundations of Economic Theory.” Philosophy and Public Affairs.
Sen, Amartya. 1985. Well-Being, Agency and Freedom: The Dewey Lectures 1984.” The Journal of Philosophy.
Shah. 2018. Coherence Arguments Do Not Entail Goal-Directed Behavior.”
Shanahan, McDonell, and Reynolds. 2023. Role Play with Large Language Models.” Nature.
Still, Sivak, Bell, et al. 2012. Thermodynamics of Prediction.” Physical Review Letters.
Strawson. 1962. “Freedom and Resentment.” Proceedings of the British Academy.
Thompson. 2007. Mind in Life: Biology, Phenomenology, and the Sciences of Mind.
Thornley. 2025. The Shutdown Problem: An AI Engineering Puzzle for Decision Theorists.” Philosophical Studies.
Turner, Smith, Shah, et al. 2021. Optimal Policies Tend To Seek Power.” In Advances in Neural Information Processing Systems.
von Neumann, and Morgenstern. 1944. Theory of Games and Economic Behavior.
Williamson. 1985. The Economic Institutions of Capitalism: Firms, Markets, Relational Contracting.
Yudkowsky. 2008. Measuring Optimization Power.”
———. 2011. Complex Value Systems in Friendly AI.” In Artificial General Intelligence (AGI 2011). Lecture Notes in Computer Science.
Zhi-Xuan, Carroll, Franklin, et al. 2025. Beyond Preferences in AI Alignment.” Philosophical Studies.

Parking lot

You should stop here. Everything elow is Claude muttering to himself while I give him research tasks.

The observer ladder

Laplace’s demon ascribes no agency.

How to evaluate a formalism-builder

The type confusion

“AI agents in practice” runs on a category mistake (Ryle 1949).

  • The LLM is not an agent as given. The weights are a conditional distribution — a generator of agents, not the sort of entity the align relation takes. The agent is generated: harness, tool loop, memory, and above all the policy for filling the context window are what turn the generator into something that acts. Simulators and role-play (janus 2022; Shanahan, McDonell, and Reynolds 2023) already make this claim at the prompt level; the 2026 product harness is the industrialised version — it pins one simulacrum, equips it with actuators, and ships it as “an agent”. This is not the boundary axis again: that axis asks where to draw the edge; this asks whether what the edge encloses is of agent type at all — a generator is not merely a small agent. (It also explains why LLM + harness is the hardest point to place on the plots: it is not one model of agency but a machine for minting them, so its coordinates smear.)
  • ”Is the LLM aligned?” is therefore type-ambiguous. Read of the instantiated system, it is an eval claim about one agent in one harness; read of the generator, it is a dispositional claim — the agents it instantiates tend to come out aligned — with entirely different success conditions. Alignment discourse slides between the readings: model cards make generator claims; incidents involve instantiated agents. Mesa-optimization (Hubinger et al. 2021) is the training-time twin of the same type split (base optimizer vs the optimizer it grows); harness-instantiation is the deployment-time version.
  • The larger system acts; it still is not a forensic agent. Deterrence needs a policy that adapts when consequences change (the Discovering-Agents test the courts have run for centuries) and a persistent identity for sanctions to reach. An instantiated agent dies at the end of its context window — sanctions do not propagate to the next instantiation except through the training loop or the deploying firm. So the forensic address resolves upward: the company (legal personhood, the civil cousin from the zoo notes) or the training pipeline, never the agent that acted. Connects to the temporal-boundary bullet below (context window as time-slice person, harness memory as prosthetic diachronic glue).
  • If promoted to the main body: the two-place opening already carries the short version in its notes; the full version is a slide of its own after “Implicit theories of agency” — the 2026-topical instance of “everyone imports a tacit model”.

From Beyond Preferences

Zhi-Xuan et al. (2025)

Angles the taxonomy forgot.

  • A third job — blueprint. Agency models are not just hired for is-work on A and ought-work about B; for the AI itself the model is a design choice. Coherence is not rationally required (Thornley 2025; Petersen 2023), coherent EU maximisation is intractable (Camara 2021), and deliberately incomplete preferences buy corrigibility — preferential gaps across contexts remove the incentive to manipulate which context the agent is in. The blueprint job wants different virtues again: coherent enough to be useful, incomplete enough to shut down.
  • Motivational currency axis (thin ↔︎ thick). Every model silently answers “what is the agent moved by”: thin preferences (bare betterness), scalar reward, thick evaluative concepts (honest, novel), reasons, role-norms. Their Evaluate–Commensurate–Decide model makes preferences constructed outputs, not primitives. Would separate models our current axes lump together (VNM and RL are both thin; forensic and role-constituted are thick).
  • The temporal boundary. Our boundary axis is synchronic; preference change (drift, volition, manipulation (Carroll et al. 2024)) and time-slice vs person-unity debates (Hedden 2015; Korsgaard 2009) raise where the agent ends in time. Live for LLM agents: is each context window a fresh time-slice person, with harness memory as prosthetic diachronic glue?
  • Role-constituted agency (in the dataframe at fog tier; parked alongside the blueprint role it exemplifies). Align to the normative ideal of a social role — the good assistant — not to preferences (Kasirzadeh and Gabriel 2023). Drift bonus: RLHF already functions as role-norm alignment (annotators judge goodness-of-a-kind) while being described as preference matching.
  • Their four preferentist theses factor along our geometry — theses 1–2 are the is/ought sides of agency-talk about A, theses 3–4 the single/multi-principal question about the B side. The paper is a book-length audit of one named point on our plot; citable validation of the two-place framing.

Economics fought these battles first

The senior discipline’s failure catalogue, mapped onto our structure.

  • Arrow–Debreu is the reassurance theorem of economics (Arrow and Debreu 1954). The welfare theorems are proved for idealised preference-maximizers and get quoted as if they described actually-existing agents — the same ought→is drift as the power-seeking theorems, opposite emotional valence. The century of battles over when equilibrium breaks (externalities, incomplete markets, information asymmetries) is the empirical record of model #1 meeting substrate.
  • Sonnenschein–Mantel–Debreu, companion to Arrow in the B-side file (Debreu 1974). Aggregation destroys structure: aggregate excess demand of perfectly rational individuals can be essentially anything. Arrow says the collective preference may not exist; SMD says even when the aggregate behaves, the theory implies nothing about it. Both are bad news for “align to humanity”.
  • Theory of the firm as boundary derivation (Coase 1937; Williamson 1985). Coase, then Williamson: the firm’s boundary is derived from transaction costs, not primitive. Challenger to the empty-quadrant claim (see the boundary × multi-agent plot note). The Coasean singularity (delegated_agent_governance) is the AI-era sequel — agent boundaries reshuffle when AI crashes transaction costs.
  • Externalities are the original alignment failures. Every agent locally aligned to its principal, aggregate outcome misaligned with everyone; gradual disempowerment is an externality story before it is anything else. Pigou vs Coase is a mechanism-design fork alignment will re-run: tax the misalignment, or assign rights and let agents bargain.
  • Coalition formation makes the partition itself strategic (coalition_games). The action question — manipulate \(v(\pi, S)\) so a preferred partition is stable — is a disempowerment mechanism: divide-and-rule played against the B side. “Humanity” is not just hard to aggregate (Arrow, SMD); it is a coalition an adversary can actively prevent from forming.
  • Meta-lesson for the talk. Alignment is speed-running economics’ arc with the same implicit agency model, and the battle reports are published.

Aligned to whom, carved how

Alignment is indexed; the verdict moves with both indices.

  • The alignment matrix (candidate slide object). Rows: dismemberments of the system (weights / product / firm / ecosystem). Columns: candidate principals (user, shareholders, society). Worked example everyone has intuitions about: Facebook’s recommender is a success story in the (product, shareholders) cell and the canonical horror story in the (product, user) and (product, society) cells — nothing about the artefact changes, we rotate the evaluation. Same exercise for Claude: the weights in-context vs the product vs the firm may each align best to a different principal. This makes the boundary axis practical: choosing where to cut is choosing which alignment claims come out true.
  • Surplus under misalignment is the norm (Jensen and Meckling 1976). Delegation persists whenever surplus exceeds monitoring costs plus bonding costs plus residual loss — nobody expects a lawyer to be aligned, only worth it net of agency costs. Reframes the AI question from purity to surplus-sharing: who captures the residual in the user–lab–model triangle? Incomplete contracting is the mature formal home (Hadfield-Menell and Hadfield 2018) — the “contract” with an AI cannot specify every contingency, which is Goodhart in contract-theory clothing. Connects to alignment_problems and delegated_agent_governance.
  • “Agent” is a homonym; the deck currently covers one sense. Agency-1: the capacity to act (the zoo’s sense). Agency-2: acting on behalf of — agency law, agency costs, fiduciary duty. Alignment discourse slides between them: “more agentic = more dangerous” is an agency-1 claim; “a good assistant is loyal” is an agency-2 claim; “AI agents” as a product category equivocates on both. Monument to the pun: the HTTP User-Agent header — the browser was named as the user’s delegate, and the ad-blocker and DRM wars are a fight over whose agent the user agent is. The role-constituted row is the agency-2 member of the zoo (a fiduciary is a role whose norms are agency-2 norms); worth saying aloud when it comes up.

Auditioning for relevance

Keep these prominent until they win or lose a slide. (Omohundro’s basic drives and the money-pump arguments won — absorbed into the is⇄ought-crossings notes.)

  • Social choice theory (Arrow 1951) — if collective agents are subject to Arrow impossibility, to what can we align? The B side again: “align to humanity” presupposes an aggregate preference that Arrow says may not exist; connects the group/institutional row to day 4. Companions from Beyond Preferences: Harsanyi’s aggregation theorem (Harsanyi 1955) is the positive case for aggregation and rests on exactly the EUT-plus-comparability assumptions the rest of the deck undermines; aggregate-preference optimisation refounds the socialist calculation debate (Hayek 1945) and is politically infeasible besides; their contractualist answer is negotiated, role-scoped norms rather than aggregated preferences.