AI persuasion, AI manipulation
Machines of loving rightthink
2023-10-05 — 2025-09-08
Wherein It Is Shown That AIs, by Lacking Out-Group Signals, Are Able to Scale Individualized Persuasion and Are Capable of Swaying Human Auditors Within Oversight and Debate Protocols
Humans, at least in the kind of experiments we are allowed to do in labs, are not very persuasive. Are machines better?
I now think: almost certainly, across a broad spectrum of persuasion tasks, yes. I was convinced by Costello, Pennycook, and Rand (2024) that AIs were already superhumanly persuasive in some settings. See also its accompanying interactive web app DebunkBot.
My reading of Costello, Pennycook, and Rand (2024)’s piece is that we tend to miss one factor when trying to understand AI persuasion: AIs can avoid participating in the identity formation and group signalling that underlies human-to-human persuasion.
There is a whole body of literature about understanding when persuasion does work in humans — for example, the work on “Deep Canvassing” had me pretty convinced that persuasion happens after the persuader has emotionally ‘got into’ the persuadee’s in-group.
“AIs can’t do that,” I thought. But I needed to realize that AIs are not in the out-group to begin with, so they don’t need to.
Aside from the patience and speed of thought, an AI also comes with the superhuman advantage of not looking like a known out-group, and maybe that is more important than looking like the in-group. I would not have picked that.
1 Field experiments in persuasive AI
We could treat field persuasion results as stress tests of scalable oversight: when participants cannot easily evaluate truth directly, models that master surface trust signals will outperform. This suggests that any governance or safety evaluation that relies on naïve human judgement will be systematically outplayed unless it incorporates calibration, interrogation rights, and group-level safeguards (Irving and Askell 2019; Bridgers et al. 2024).
For example, Costello, Pennycook, and Rand (2024). This is pitched as AI-augmented de-radicalization, but you could change the goal and think of it as a test case for AI-augmented persuasion in general. Interestingly, the AI doesn’t seem to need to resort to traditional workarounds for human tribal reasoning, such as deep canvassing. The potential to scale up individualised mass persuasion could become the dominant game in town, at least until some new equilibrium of persuasive chatbots has been reached.
See also Salvi et al. (2024), Schoenegger et al. (2025), and the follow-up Kowal et al. (2025).
2 As reward hacking
See human reward hacking.
3 In debate and oversight protocols
See also Debate update: Obfuscated arguments problem
Give evaluators the right to interrogate model claims. Humans rarely ask maximally informative questions by default; prompting or tooling can improve query quality and reduce overreliance (Rothe, Lake, and Gureckis 2018; Bridgers et al. 2024).
4 Collective Deliberation vs. Coordinated Persuasion
Habermas’s evil twin.
- In oversight experiments, structured discussion can improve collective accuracy via information pooling and error cancellation (Bergman et al. 2024; Navajas et al. 2018; Bahrami et al. 2010).
- In persuasion settings, the same coordination channels create cascade risks: correlated updates, herding, and fast belief shifts when messages are optimized for group-level influence.
Deliberation likely helps only when
- message provenance is auditable,
- dissent is protected,
- timing/randomization breaks echo chambers, and
- confidence claims are penalized when later falsified (Irving and Askell 2019).
I am curious what this does to Civic tech.
5 Tipping points in mass persuasion
TBD
- Does persuasion lower the activation energy for cascades (Barnes and Christiano 2020; Irving, Christiano, and Amodei 2018; Michael et al. 2023).
6 Spamularity and the scam economy
See Spamularity.
7 Persuasion as a Control Problem
If powerful models are subject to control protocols (monitoring, containment, restricted affordances), the human-in-the-loop becomes a primary attack surface. Social engineering is a known vector for bypassing security, and the same applies to AI oversight. An escaping or scheming system doesn’t need root access if it can persuade the auditor to approve a risky action, overlook an anomaly, or relax a safeguard (Korbak et al. 2025), cf Shlegeris and Greenblatt.
8 The future wherein it is immoral to deny someone the exquisite companionship and attentive understanding of their own personal AI companion
See artificial intimacy.
9 Incoming
AI systems out-persuade expert humans. Panic? reviews Hackenburg et al. (2026)
AI Agents and Democratic Resilience | Knight First Amendment Institute
Artificial Intelligence and Democratic Freedoms | Knight First Amendment Institute
-
Think you understand how social media really works? Join Australia’s newest and most innovative AI-based cybersecurity competition!
Through a controlled, artificial social media landscape, competitors will be charged with the manipulation of social media narratives by developing and deploying bots that amplify and suppress target messages.
Capture the Narrative provides interactive insights into how social media platforms can be manipulated, raising awareness of the potential for abuse and the importance of digital cyberliteracy.
The top team of 1-4 students will take home AU$5,000 and our competition is open to students from every Australian University.
Beth Barnes on Risks from AI persuasion
Dynomight, in I guess I was wrong about AI persuasion, emphasizes some different things that I was thinking of. See also the comments
Reward hacking behavior can generalize across tasks — AI Alignment Forum
-
an easy-to-use integration of leading AI models (from openai, anthropic, meta, google, and more) into qualtrics, allowing researchers to design experiments where humans and AI interact with just a few clicks. made by researchers at mit (@hauselin, @tawabsafi, @thcostello).
test a new intervention, create a semi-structured interview, or build something entirely new. we’re trusted by leading scientists at mit, stanford, university of oxford, london school of economics, university college london, and the university of toronto.
vegapunk takes minutes to set up, instantly scales to thousands of participants, is safe and secure, and highly customizable. see documentation here

