Should our journal publish AI-drafted manuscripts?
Forget both truth and beauty, I want to know about opportunity costs
2026-08-24 —
2026-09-13
quality 5.4
In Which a Minimal Economic Model of the Alignment Journal Is Utilized, Wherein Authorship Types Are Bifurcated Into High-Quality and Slop, and Submission Fees Are Salvaged as a Congestion Charge.
academe
AI safety
collective knowledge
doing internet
economics
faster pussycat
how do science
incentive mechanisms
innovation
institutions
machine learning
mind
provenance
sociology
technology
Status: Draft model, AI-assisted, not thoroughly checked.
This is a purely personal exploration of a policy problem and definitely does not represent the opinion of the Alignment Journal itself. Given the context, I suppose I had best disclose my own AI usage in this article: transformative. Although the original model was mine, it was made way better by iterative refinement and re-drafting by AI, and by no means would I have had time to write it by hand.
Figure 1
At the Alignment Journal we have been debating whether to accept AI-drafted manuscripts for review. This is a relatively high-leverage question, as detecting AI-drafted prose is (currently) surprisingly feasible, making it a cheap signal for us to use in desk review to filter out low-quality papers.1
The perfect policy would optimize for the overall quality of the journal’s output with regard to how that serves our readers. On one hand, AI drafting lowers the cost of producing a manuscript, which may increase the flow of good work into our inboxes. On the other hand, it might increase the flow of low-quality AI slop into those inboxes. That is, we worry about AI slop spam, asaremanyothers(Gartenberg et al. 2026).
Insofar as the base costs of a high-quality paper exceed those of a slop manuscript, we might feel that permitting AI drafting provides a relatively larger benefit to slop authors than to substantive authors.
So, while we can detect AI drafting, should we do so? And should we use it to desk-reject papers?
Here is a minimal economic model of that question, in the vein of Agrawal, Gans, and Goldfarb (2019), which models AI as a fall in the price of prediction. I have made stylized assumptions to make the model tractable before it is realistic or empirically calibrated. Where I need to choose some numbers, I have taken inspiration from the Alignment Journal, since that is closest to my heart.
Here, I assume AI drafting reduces the cost of producing prose: the writing-up stage of science gets cheap while the substance stage, for now, does not (as per Kwa et al. 2025). That is a shaky assumption even now (for example, both my drafting and substantive work in writing this very post were AI-expedited) and it will get shakier as AI progresses.
My goal here is not to persuade anyone of a particular optimal policy, but to lay out a concrete model of the first- and second-order economic effects of AI drafting so that we can debate these things better. In particular, I think that too often AI drafting discussions exclude the economic and strategic effects on the production of scientific knowledge. Instead we tend to discuss the problem in purely moral terms (AI is “efficient” or “lazy”), or aesthetic terms (“AI prose is hard to read”), which are important but incomplete. Our duties in publishing are not only to truth, beauty and justice, but also surely include the economic and strategic management of scientific knowledge production. I think all these are useful lenses, to be clear, but the models of economic effects seem undersupplied, and so here we find ourselves.
AI-drafting preferentially incentivises production of poor papers, which crowd out good ones.
A ban enforced by an AI-drafting detector passes through three regimes as the detector’s true-positive rate rises.
At a low true-positive rate it changes nothing: it desk-rejects the same fraction of both high- and low-quality papers.
At a middling true-positive rate it deters the good researchers from drafting with AI, who go back to hand-drafting as an expensive signal of their quality, while slop authors keep drafting with AI and lose a fraction of their manuscripts to desk rejection. (Good authors give up AI before slop authors do, at every capability level.)
At a high true-positive rate slop authors give up AI too, everyone goes back to hand-drafting, and the no-AI outcome returns.
Alternatively we could just charge all authors submission fees, which lets everyone use AI while disincentivising spam, potentially even better than an automated ban. The optimal fee is ballpark USD 5,000 per submission, which is outside the Overton window and moreover impractical for most authors to realistically pay. No-one is proposing this at the moment, not least because it is a business model indistinguishable from that of predatory journals.
A hybrid, where authors may pay to skip the slop detector, can inherit the virtues of both a ban and a fee. The optimal fee comes to about USD 7,000 per AI-drafted submission. No-one is proposing this either, for similar reasons to the previous.
Regarded purely as an instrument for controlling slop rates, the classifier-based ban is a clumsy approximation to a congestion charge; the hybrid version makes that charge explicit.
I have not modelled various other factors such as reputational effects, which could significantly alter the incentives for both high-quality and slop authors.
1 Setup
I assume there are two types of authors: high-quality authors (type \(H\)) and slop authors (type \(L\)). They differ only in the substance of their manuscripts and behave strategically about the presentation. The proportion of high-quality authors is \(\lambda\), and the rest are slop authors. Articles have apparent substance\(s\), which is the quality the manuscript presents to a reviewer—the quality a reviewer can assess. I measure it in citoms, my idiosyncratic unit of “citability,” which happens to be what citations count. These are themselves an imperfect proxy for our true desired output, scientific progress\(v\), which is what a reader gets if the manuscript is published. I measure that in scientillas.2
A high-quality author spends substance time\(t\) on each manuscript—the hours of research and thinking as opposed to writing up. The result is a manuscript that reads as good as it actually is, with as many citoms as scientillas:
\[
v_H =s_H = \theta \sqrt{t},
\]
where \(\theta\) is the author’s skill, which I put in the formulas to mark where a spread of skill would enter … but actually I ignore it and normalize to \(\theta = 1\) throughout.
The rest of the authors produce slop, which in my model means manuscripts worth \(s_L\) citoms at a glance but zero scientillas upon deep reading, \(v_L = 0\), made with the minimum substance time \(t_{\min}\). All of which is to say, the \(H\) authors aren’t deceptive, whereas the \(L\) authors produce low-quality outputs masquerading as adequate.
Every manuscript additionally requires drafting time \(d\): hand-drafting costs \(f\), AI drafting costs \(\delta = f/a\), where \(a\) is the AI’s drafting efficiency. An author has a time budget \(T\) per period (year?), so an author who produces \(n\) manuscripts at a cost of substance effort time \(t\) satisfies \(n(t + d) = T\). We write \(t_H\) and \(n_H\) for the substance time and manuscript count an \(H\)-type author chooses, and \(n_L\) for a slop author’s manuscript count; slop authors have no substance time to choose.
Reviewing a manuscript to the journal’s standard takes \(h\) reviewer-hours, which we call the depth of review, and there are \(R\) total reviewer-hours, so the journal can review at most \(M = R/h\) manuscripts. Facing \(N\) submissions, it reviews \(\min(N, M)\) of them and rejects the rest unread, so a submission’s chance of being read at all is the coverage rate
\[
r = \min\!\left(1, \tfrac{M}{N}\right).
\]
In this simplified model, the only desk-reject stage is the AI-drafting classifier of Policy B, which sits in front of triage and costs no reviewer time; the free-for-all has none. A manuscript that receives a review is accepted with a probability depending on both its substance \(s\) and the review time \(h\),
i.e., deeper reviews detect substance more precisely. Here \(\bar{s}\) is the journal’s bar, the number of citoms at which a reviewed manuscript is accepted half the time. We write \(\Lambda_H\) and \(\Lambda_L\) for the acceptance probability of a reviewed high-quality manuscript and a reviewed slop manuscript, respectively. The journal holds the depth \(h\) fixed with respect to queue size.
Authors are assumed to be careerist, in that they care only about a payoff \(b\) per accepted paper. The payoff is earned by clearing the bar in citoms; scientillas earn an author nothing beyond that. A high-quality author facing a drafting cost \(d\) thus solves
maximizing accepted papers per unit of time. No author can move the coverage rate \(r\), and it scales the payoff of every candidate \(t\) identically, so it drops out of the optimization. As such, congestion never influences the choice of how much effort to dedicate to substance. What does influence the choice is the drafting cost \(d\), which is \(f\) by hand or \(\delta = f/a\) with AI. Every hour spent on the substance of one manuscript is an hour not spent starting the next, and \(d\) controls how much that next manuscript would have cost. The first-order condition is
As AI capability rises, \(\delta\) falls, the next manuscript gets cheaper, and the optimum tilts toward more, thinner papers.
Now let us consider the target audience, the readers. A journal’s goal, as a benevolent planner, is to maximize the readers’ welfare; we might describe this in two ways that end up being quite similar.
Firstly, we might consider the average scientillas per published article,
\[
W \;=\; \mathbb{E}\left[\, v \mid \text{accepted} \,\right].
\]
Recall that a high-quality manuscript delivers \(s_H\) scientillas, while slop delivers none even though it presents \(s_L\) citoms. This metric is something like the average amount of science I will find sampling randomly from the published literature.
Secondly, we might consider the number of good papers published per period, the total amount of science we are pumping out to all readers. Of the \(\lambda\,n_H\) high-quality manuscripts written, a fraction \(r\) get reviewed and a fraction \(\Lambda_H\) of those pass, so the flow of high-quality papers into print is the throughput
\[
Q \;=\; r\,\lambda\,n_H\,\Lambda_H.
\]
Either seems a reasonable objective for the journal. Happily, every ranking of one policy against another below comes out the same way under both, so we can track the pair without choosing between them; they disagree only about how high to set the fee in Policy D. Both count scientillas, which measure public-good outcomes. The authors’ incentives by contrast, count citoms, career-wise goals measured in citations. Slop writing is the gap between these two.
I make three other simplifying assumptions:
AI drafting here introduces no errors and incurs no cognitive debt3
Accepted slop delivers zero value but is not actively harmful.4
The review budget is exogenous: \(R\) is money divided by the going wage for qualified reviewer attention, and the model holds both fixed.
Code
import numpy as npimport matplotlib.pyplot as pltfrom livingthing.matplotlib_style import set_livingthing_styleset_livingthing_style()P =dict(theta=1.0, s_L=0.30, sbar=0.50, t_min=0.05, f=0.50, T=1.0, lam=0.25, R=16.0, h=10.0, b=1.0)T_GRID = np.geomspace(P["t_min"], 25.0, 4000)def logistic(x):return1.0/ (1.0+ np.exp(-np.clip(x, -60, 60)))def accept_H(t, p=P):return logistic(p["h"] * (p["theta"] * np.sqrt(t) - p["sbar"]))def accept_L(p=P):return logistic(p["h"] * (p["s_L"] - p["sbar"]))def best_H(d, p=P, fee=0.0):"""Substantive author's optimum: maximize (b*accept - fee)/(t+d). The coverage rate r is a common factor, so it never appears here. Returns (value per unit time, t*).""" obj = (p["b"] * accept_H(T_GRID, p) - fee) / (T_GRID + d) i =int(np.argmax(obj))return obj[i], T_GRID[i]def outcomes(nH, wH, tH, nL, wL, p=P):"""Coverage, mean substance per article, and good-paper flow.""" N = p["lam"] * nH * wH + (1- p["lam"]) * nL * wL r =min(1.0, (p["R"] / p["h"]) / N) if N >0else1.0 aH = p["lam"] * nH * wH * accept_H(tH, p) aL = (1- p["lam"]) * nL * wL * accept_L(p) sH = p["theta"] * np.sqrt(tH) W = aH * sH / (aH + aL) if aH + aL >0else0.0returndict(r=r, W=W, thruH=r * aH, tH=tH, sH=sH)def solve_A(a, p=P):"""Free-for-all: everyone AI-drafts, triage to capacity.""" delta = p["f"] / a _, tH = best_H(delta, p)return outcomes(p["T"] / (tH + delta), 1.0, tH, p["T"] / (p["t_min"] + delta), 1.0, p)def mu_star(a, p=P):"""True-positive rate above which slop producers abandon AI drafting."""return1.0- (p["t_min"] + p["f"] / a) / (p["t_min"] + p["f"])def mu_sub(a, p=P):"""True-positive rate above which high-quality authors abandon AI drafting.""" Vh, _ = best_H(p["f"], p) Va, _ = best_H(p["f"] / a, p)return1.0- Vh / Vadef solve_B(a, mu, p=P):"""Classifier ban: AI-drafted papers desk-rejected with true-positive rate mu.""" delta = p["f"] / a dH, wH = (delta, 1- mu) if mu < mu_sub(a, p) else (p["f"], 1.0) dL, wL = (delta, 1- mu) if mu < mu_star(a, p) else (p["f"], 1.0) _, tH = best_H(dH, p)return outcomes(p["T"] / (tH + dH), wH, tH, p["T"] / (p["t_min"] + dL), wL, p)def solve_C(a, p=P, eps=1e-3):"""Fee at slop's uncongested break-even: slop exits entirely.""" delta = p["f"] / a fee = p["b"] * accept_L(p) + eps _, tH = best_H(delta, p, fee=fee) out = outcomes(p["T"] / (tH + delta), 1.0, tH, 0.0, 0.0, p) out["fee"] = feereturn out
2 Policy A — Free-for-all
Under this polity anyone may use AI, and we triage submissions according to capacity (if there are too many papers we reject the excess). The equilibrium is easy to compute: the substance choice depends only on the drafting cost \(\delta = f/a\), submission counts follow from the time budget, and coverage follows from the counts. tl;dr Under this policy everyone drafts their manuscripts with AI, as it is cheaper and goes unpunished.
Code
a_grid = np.geomspace(1, 100, 25)eqA = [solve_A(a) for a in a_grid]fig, axes = plt.subplots(1, 3, figsize=(9.5, 3), sharex=True)axes[0].plot(a_grid, [q["tH"] for q in eqA])axes[0].set_title("substance time $t_H$")axes[1].plot(a_grid, [q["r"] for q in eqA])axes[1].set_title("coverage $r$")axes[2].plot(a_grid, [q["W"] for q in eqA], label="scientillas per article $W$")axes[2].plot(a_grid, [q["thruH"] for q in eqA], ls="--", label="good papers per author-period $Q$")axes[2].set_title("welfare")axes[2].legend(fontsize=8)for ax in axes: ax.set_xscale("log") ax.set_xlabel("AI capability $a$")fig.tight_layout()plt.show()
Figure 2: The free-for-all as AI drafting capability \(a\) grows: substance time per high-quality manuscript, the coverage rate \(r\), and the two welfare metrics. \(t_H\) falls somewhat; the large damage comes from crowding.
Optimistically, we might hope that cheaper AI drafting frees up time compared to the manual alternative, allowing for more substance per paper. That does not happen under this model. The left edge of every capability sweep, \(a = 1\), is where AI drafting saves nothing, so everyone might as well hand-draft; I call that no-AI state the hand-drafting equilibrium. For the high-quality tier \(H\), the substance time per manuscript \(t_H\) falls modestly as \(a\) rises, from \(0.46\) to about \(0.32\): cheaper drafting makes each manuscript cheaper, so authors choose quantity over quality. The fall is not too precipitous though: high-quality manuscripts stay above the bar (\(s_H\) drifts from \(0.68\) down to \(0.56\) scientillas, against a bar of \(0.5\) citoms), and their authors write three of them where they used to write one. For the low-quality tier \(L\), the slop authors’ output grows much faster, because a slop manuscript is nearly all drafting time: \(n_L\) rises from under \(2\) to over \(18\).
The downside is visible through the effect on the queue, albeit differently in each of the welfare metrics. Scientillas per article: the submission pool ends up \(95\%\) slop, slop passes a full review \(12\%\) of the time, so most of what the journal accepts is slop, and \(W\) falls by two thirds. Throughput: coverage falls from roughly \(1\) to \(0.11\), so the journal publishes about four times fewer high-quality papers — even though the high-quality authors are writing more of them than before. Most of their extra manuscripts are simply never read. In the units of the model that is \(Q\) falling from \(0.22\) to \(0.06\) good papers per author-period: a journal drawing on a hundred authors goes from printing 22 good papers a period to 6. Total output shrinks as well, from \(0.38\) to \(0.24\) papers per author-period, slop included: coverage collapses so far that even slop acceptances fall.
The collapse is smooth and the equilibrium is unique. The coverage rate cancels out of every author’s problem, so there is no feedback loop to amplify anything, and no tipping point. The Diluted review appendix describes a review technology under which there is one.
3 Policy B — Automated ban on AI-drafted manuscripts
A stylometric classifier (e.g. Pangram) flags manuscripts it thinks are AI-drafted, and we desk-reject those.5 This is the only desk-reject stage in the model; everything it passes goes into the capacity triage of the Setup. We assume the classifier has true-positive rate \(\mu\), which the machine-learning literature calls its recall: it correctly flags a fraction \(\mu\) of AI-drafted manuscripts. For now, I assume it has a zero false-positive rate to keep it simple.6 Desk-rejected papers consume no reviewer time. The rest get a normal review.
Now, knowing this policy, everyone chooses how to write their papers. For low-quality authors, the trade-off is between the time saved by using AI and the risk of being desk-rejected. Their acceptance probability if they get past desk-rejection is the same however the manuscript was drafted, so it cancels from the comparison, and what remains is: papers per hour with AI, discounted by the survival rate \((1-\mu)\), against papers per hour by hand. They keep AI drafting while \(\mu < \mu^*(a)\), where
In this setting, the classifier’s true-positive rate is pivotal. The true-positive rate needed to make slop authors give up AI depends only on the time costs of producing slop, and it rises with AI capability toward a limit given by the share of a hand-drafted slop manuscript’s time that goes on drafting, \(f/(t_{\min}+f)\). At my parameters, that limit is about \(0.91\), so no classifier below it can ever deter slop authors from AI drafting, however capable the AI becomes. Whether real classifiers reach that is contested: an independent test of six deployed detectors found accuracies between \(55\%\) and \(97\%\) on ordinary, non-adversarial text (Akram 2023), while the strongest commercial classifier, Pangram, reports a false-negative rate of \(0.34\%\) at a false-positive rate of \(0.004\%\) on unmodified LLM output, comfortably above any threshold in this model (Glickenhaus et al. 2026), so it seems credible that Pangram is useful for actually existing off-the-shelf LLM drafting.
On the other hand, we should leaven our enthusiasm with caution: the model needs a true-positive rate sustained against slop authors who are actively rewriting to evade detection, because once a detector cuts into their earnings, defeating it becomes worth their effort. The arms race is taking place in the real world and the current advantage of AI detectors may not be sustainable. Pangram’s own report puts its true-positive rate on AI-edited human text at only \(55\)–\(65\%\)(Glickenhaus et al. 2026). Independent tests on academic abstracts find a true-positive rate of about \(95\%\) on unmodified output, falling to \(24\%\) once rewrite prompts are searched against the detector (Ren, Raghavan, and Garg 2026) and below \(4\%\) after a commercial humanizer (Karr et al. 2026), both well under the limit of \(\mu^*\). Nonetheless, against unmotivated slop, it can be exploited for now. Note, however, that we expect the threshold to move. Because \(\mu^*(a)\) rises with capability, a ban that deters slop today can eventually fail without any policy change and without any decline in the detector. Hold \(\mu\) fixed, let \(a\) grow, and the ban passes from deterring slop, to taxing only the compliant, to doing nothing see 4.
The high-quality authors have a different behaviour threshold. Write \(V(d)\) for the best rate of expected acceptances per unit time that a high-quality author can achieve at drafting cost \(d\). High-quality authors abandon AI drafting once \((1-\mu)\,V(\delta)\) falls below \(V(f)\), that is at
\[
\mu_H(a) = 1 - \frac{V(f)}{V(\delta)}.
\]
The two thresholds are strictly ordered: \(\mu_H(a) < \mu^*(a)\) for every \(a > 1\) (at \(a = 1\) both are zero). To see this, we evaluate the hand-drafting value at the AI-drafting optimum \(t^*_\delta\), which gives \(V(f) \ge V(\delta)\,\frac{t^*_\delta + \delta}{t^*_\delta + f}\), and the ratio \(\frac{t + \delta}{t + f}\) increases in \(t\). In other words: the more substance time we invest in a manuscript, the smaller the share of its cost that AI drafting saves, so the less detection risk an author will accept to keep using it. A detector therefore stops high-quality authors from using AI strictly before it stops slop authors. Neither type is ever deterred from submitting: at worst, slop authors revert to hand-drafting, which restores the hand-drafting equilibrium’s spam rate, never less. Being flagged costs a high-quality author a manuscript full of work; it costs a slop author almost nothing.
Code
mus = np.linspace(0.02, 0.98, 120)W_pre = solve_A(1.0)["W"]fig, ax = plt.subplots(figsize=(6.5, 3.5))for a, col inzip((2, 5, 15, 100), ("#b5541c", "#7f4c94", "#2d6e8e", "#3a7d44")): ax.plot(mus, [solve_B(a, mu)["W"] for mu in mus], color=col, label=f"$a = {a}$") ax.axvline(mu_sub(a), color=col, ls="--", lw=0.8) ax.axvline(mu_star(a), color=col, ls="-", lw=0.8, alpha=0.5)ax.axhline(W_pre, color="#555555", ls=":", lw=1)ax.set_xlabel(r"classifier true-positive rate $\mu$")ax.set_ylabel("scientillas per article $W$")ax.legend()plt.show()
Figure 3: Scientillas per published article, \(W\), under the classifier ban, against true-positive rate \(\mu\), at several capability levels. Dotted horizontal: the hand-drafting equilibrium. Dashed verticals: \(\mu_H\), where high-quality authors give up AI drafting. Solid verticals: \(\mu^*\), where slop authors do. Above \(\mu^*\) the hand-drafting equilibrium returns.
We observe three regimes. Below \(\mu_H\), the ban deters nobody and changes nothing. Both grades of author continue drafting with AI, so the classifier desk-rejects the same fraction \(\mu\) of each type’s manuscripts. The scientillas per accepted article remain untouched, and the review slots freed by desk rejection compensate for the good manuscripts that were desk-rejected, so both metrics stay at their free-for-all values. Between \(\mu_H\) and \(\mu^*\), the ban’s only behavioural effect is on the wrong people: high-quality authors have gone back to hand-drafting, writing one manuscript where they could have drafted three; slop authors still submit at their full rate and pay the cost of having a fraction \(\mu\) of their manuscripts desk-rejected. Welfare rises across this range of true-positive rates only because \((1-\mu)\) shrinks the surviving spam, not because anyone has stopped. The upward step in \(W\) at \(\mu_H\) occurs when high-quality authors switch to hand-drafting. By the definition of \(\mu_H\), their rate of accepted papers remains unchanged at the switch, so \(Q\) does not move; what changes is that each hand-drafted manuscript carries more substance (\(s_H\) rises from \(0.56\) to \(0.68\) at \(a = 100\)), and \(W\) rises by exactly that ratio. The ban’s one benefit in this range, on this metric, is forcing good authors to write fewer, fuller papers. Above \(\mu^*(a)\), everyone hand-drafts and the hand-drafting equilibrium reappears.
Figure 3 holds capability fixed and sweeps the detector; Figure 4 holds the detector fixed and lets capability grow, in the same panels as Figure 2, which is the direction the world actually moves in. The same three regimes appear, now read from right to left in \(a\): the ban deters slop while \(\mu^*(a)\) is still below the detector’s rate, taxes only the compliant once \(\mu^*(a)\) has climbed past it, and does nothing once \(\mu_H(a)\) has too.
Code
mu_fix =0.5eqB = [solve_B(a, mu_fix) for a in a_grid]a_star = P["f"] / ((1- mu_fix) * (P["t_min"] + P["f"]) - P["t_min"])lo, hi =1.0, 1e4for _ inrange(60): mid = np.sqrt(lo * hi) lo, hi = (mid, hi) if mu_sub(mid) < mu_fix else (lo, mid)a_H = np.sqrt(lo * hi)fig, axes = plt.subplots(1, 3, figsize=(9.5, 3), sharex=True)grey =dict(color="#888888", alpha=0.7)axes[0].plot(a_grid, [q["tH"] for q in eqA], **grey)axes[0].plot(a_grid, [q["tH"] for q in eqB])axes[0].set_title("substance time $t_H$")axes[1].plot(a_grid, [q["r"] for q in eqA], **grey)axes[1].plot(a_grid, [q["r"] for q in eqB])axes[1].set_title("coverage $r$")axes[2].plot(a_grid, [q["W"] for q in eqA], **grey)axes[2].plot(a_grid, [q["thruH"] for q in eqA], ls="--", **grey)axes[2].plot(a_grid, [q["W"] for q in eqB], label="scientillas per article $W$")axes[2].plot(a_grid, [q["thruH"] for q in eqB], ls="--", label="good papers per author-period $Q$")axes[2].plot([], [], label="free-for-all", **grey)axes[2].set_title("welfare")axes[2].legend(fontsize=8)for ax in axes: ax.set_xscale("log") ax.set_xlabel("AI capability $a$") ax.axvline(a_star, color="#555555", lw=0.8) ax.axvline(a_H, color="#555555", lw=0.8, ls="--")fig.tight_layout()plt.show()
Figure 4: The classifier ban at a fixed (low) true-positive rate \(\mu = 0.5\) as AI capability \(a\) grows, in the panels of Figure 2, with the free-for-all in grey for comparison. Vertical lines: where \(\mu^*(a)\) passes \(0.5\) (left) and where \(\mu_H(a)\) does (right). Left of the first line the ban deters slop authors from AI drafting, so everyone hand-drafts and the hand-drafting equilibrium of \(a = 1\) persists whatever \(a\) is; between the lines it deters only the high-quality authors, who hand-draft while slop authors keep using AI; right of the second it deters nobody and the free-for-all returns.
The standard objection to a ban is that it takes the time savings of AI drafting away from the compliant. My model agrees that the ban forfeits those savings, but so does the free-for-all, which converts them into extra manuscripts that mostly go unread. The savings are only worth having under a policy that can turn them into published papers, and that is what the submission fee of Policy C does.
NB: The assumed zero false-positive rate is important; I left it untouched for simplicity, but since it punishes compliant authors, even a small false-positive rate can significantly alter the outcomes near \(\mu^*\) economically, not to mention erode trust.
4 Policy C — submission fees
Under this policy, we charge \(P\) per submission, allow anything, triage and review whoever pays for it. To be clear: this is not on the table for the Alignment Journal, but is included as a baseline.
An economist would say that submission fees are always charged, in waiting time if not in cash, and a journal chooses some mix of the two whether or not it admits to pricing (Cotton 2013). This does not change how it feels to be charged such a fee, nor the fact that many authors have no cash budget to pay it. Nonetheless…
Recall that a submission’s expected private return is \(b\) times its acceptance probability: \(r\,b\,\Lambda_L\) for slop, and roughly \(r\,b\,\Lambda_H\) — much larger — for a high-quality paper. Any fee between those two numbers is separating: submitting slop now loses money, while submitting good work still pays. Nobody is forced back to hand-drafting, because the fee is the same however the paper was drafted — which is fine, since here at least we don’t care about provenance except as a proxy for spamminess.
If the submission fee is too low, slop authors keep entering until congestion drives their expected return down to the fee, \(r\,b\,\Lambda_L = P\). That pins the coverage rate at \(r = P/(b\,\Lambda_L)\): slop manuscripts fill every review slot that high-quality manuscripts do not. Raising \(P\) buys back coverage, but the accepted papers stay mostly slop until the fee exceeds what a slop author values each submission at.
Full exclusion of slop is sadly expensive. A slop author facing an uncongested queue (\(r = 1\)) expects \(b\,\Lambda_L\) per submission, so the fee at which slop authors stop submitting entirely is
about an eighth of the private value of an acceptance (at my parameters). \(P^*\) does not depend on \(a\): the review depth \(h\) is fixed, so a slop manuscript’s acceptance odds are the same however large the flood, and the fee has to do all the deterring by itself. On the other hand, this gives this policy the nifty feature of being able to fully exclude slop authors with a single fee setting even as capabilities advance, unlike, say, the ban on AI-generated content, which has to be re-checked as \(\mu^*\) or \(a\) changes. At \(P^*\) every slop author stops submitting, at least if they are rational expected-value maximizers.
Code
fig, (ax, axt) = plt.subplots(1, 2, figsize=(9.5, 3.4))eqC = [solve_C(a) for a in a_grid]for ax_, key in ((ax, "W"), (axt, "thruH")): ax_.plot(a_grid, [q[key] for q in eqA], color="#b5541c", label="A: free-for-all")for mu, ls in ((0.6, "--"), (0.95, "-")): ax_.plot(a_grid, [solve_B(a, mu)[key] for a in a_grid], ls, color="#2d6e8e", label=f"B: classifier, $\\mu={mu}$") ax_.plot(a_grid, [q[key] for q in eqC], color="#3a7d44", label="C: fee") ax_.set_xscale("log") ax_.set_xlabel("AI capability $a$")ax.set_ylabel("scientillas per article $W$")axt.set_ylabel("good papers per author-period $Q$")ax.legend(fontsize=8)fig.tight_layout()plt.show()
Figure 5: The three policies as AI capability grows. Classifier at true-positive rate \(0.6\) and \(0.95\); fee at the full-exclusion level \(P^*\). Left: scientillas per published article \(W\). Right: good papers published per author-period \(Q\). Only the fee regime turns more capable AI into more good papers published.
Interestingly, under the fee policy, more AI capability is simply good. High-quality authors keep all the time AI saves them and spend it writing more manuscripts. Huzzah! The journal, its queue protected, has the capacity to review and publish them; and the flow of good papers into print more than doubles across the plotted range. Under the other regimes, we burn that time surplus: the ban burns it by forcing high-quality researchers back to hand-drafting, and the free-for-all converts it into unread submissions. The fee curve in the figure is computed at \(P=P^*\), the price at which slop authors stop submitting, so no slop reaches the accepted papers at all. At any lower fee, slop authors keep submitting until the queue congestion lowers their expected return to the fee, so slop manuscripts fill every review slot the high-quality manuscripts leave free, and reviewers still accept some of those.
The size of the fee’s lead in the left panel has a simple decomposition: \(W\) is the share of published articles that are high-quality, times the scientillas \(s_H\) in a high-quality one. We can understand the gap as a mismatch between the proxy (was AI used to write this?) and the true target (was this article high-quality?). The best the ban can do, at any true-positive rate, is restore the hand-drafting equilibrium, which at these parameters means about 58% high-quality articles. A sufficiently large fee removes all slop from the queue, leading to a situation where the high-quality share is \(100\%\). A 58% high-quality share times 0.68 scientillas per good article under the ban (=0.4 scientillas per article), against a 100% share times about 0.60 (=0.6 scientillas per article) under the fee at \(a = 100\).
An entailment of that calibration is that the journal was hypothetically going to be about \(40\%\) bullshit in the absence of AI drafting. That number follows from parameter choices I made for \(\lambda\), \(s_L\) and \(h\), which, as I said, are not crazy. Readers who believe their journal is purer or better than that should raise \(\lambda\) or \(h\) and watch the fee’s lead over the ban shrink accordingly. The fee’s advantage on this metric arises from the amount of hand-drafted slop it eliminates, so the dirtier we think journals already are, the stronger the case for pricing submission over policing provenance, and vice versa.
Of course, many journals cannot charge submission fees for reasons of custom and equity, and fear that no matter how eloquent the justification it will still look like predatory journal behaviour. The Alignment Journal is not exempt from such reasonings. In such cases, the deadweight losses of dealing with spam are still imposed on someone, and in fact authors bear them, paying in lost time queuing for a review slot: with cash fees near zero, first-response delay is the de facto submission price (Azar 2005).
As such Policy C is useful as a benchmark. It charges for review slots in cash rather than in waiting time, the economists’ platonic ideal. The gap between it and the best policy we can feasibly adopt is the price we pay for dwelling in the earthly realm of imperfect expected utility maximizers and departmental budget constraints.
If a journal did wish to explore this fee-charging route, it would need to do it carefully and transparently. For instance, a journal that pays its reviewers, as the Alignment Journal does, could spend the fee revenue on more reviewing, so the fee would partly fund its own capacity. We would need to do something to allay the fears of predatory behaviour (for example, compromising review quality to accept more submissions with less review quality), etc.
Another option is a fee refunded on acceptance. In expectation it charges bad papers more than good ones, which seems fairer to the authors of good papers. This is not free from moral hazard either.
Elaborations are left as homework.
5 Policy D — authors pay to skip the classifier
Authors could pay a fee \(P\) per manuscript to have it reviewed without passing through the classifier. Manuscripts that do not pay go through Policy B: the classifier runs, and what it flags is desk-rejected. Hand-drafted manuscripts are never flagged, so their authors have no reason to pay. The fee is only ever paid by authors who drafted with AI and would rather not gamble on the detector. It imposes Policy C’s submission fee, but charges only AI drafters. Each manuscript now has three routes: draft by hand and pay nothing, draft with AI and risk desk rejection, or draft with AI and pay the fee.
Consider the slop authors first. A slop manuscript’s acceptance odds do not depend on how it was drafted, so if it reaches review it is worth \(r\,b\,\Lambda_L\) to its author however it got there. An author who drafts with AI and does not pay has a fraction \(\mu\) of their manuscripts desk-rejected, which costs them \(\mu\, r\,b\,\Lambda_L\) per manuscript in expectation. So paying the fee beats risking desk rejection only when the fee is less than \(\mu\, r\,b\,\Lambda_L\). Paying the fee beats drafting by hand only when the fee is less than \(\mu^*\, r\,b\,\Lambda_L\), by the same arithmetic that gave us \(\mu^*\) under the ban. So a slop author pays the fee only when
\[
P \;<\; \min(\mu, \mu^*)\; r\,b\,\Lambda_L.
\]
Call the right-hand side the floor fee, \(\underline{P} = \min(\mu, \mu^*)\, r\,b\,\Lambda_L\): below that price every slop author pays to skip the detector, and above it none does. Any fee the journal wants slop authors not to pay must at least that high. At a fee above the floor, slop authors ignore the exemption option and take their chances with the classifier as under Policy B, to wit: below price \(\mu^*\) they draft with AI and submit anyway, writing off the fraction \(\mu\) of their manuscripts that is desk-rejected as a cost of doing business; above \(\mu^*\) they draft by hand, are never flagged, and submit at the hand-drafting equilibrium rate. The floor is lower than Policy C’s full-exclusion fee \(P^*\) by the factor \(\min(\mu, \mu^*)\,r\), because skipping the classifier is only worth what the classifier would have cost the author, and in a congested queue a submission is reviewed only with probability \(r\), which scales its expected return—and so in turn the fee an author will pay—by the same factor (\(P^*\) was defined at \(r = 1\)). That is the floor, not the fee the journal ends up charging: the optimal exemption fee comes out above\(P^*\) at my parameters, as we see in a moment, because it is set by congestion among the high-quality authors rather than at the rate takes to deter slop.
which is a fancified version of the same calculation in Policy C, differing in a couple of elaborations.
First, the coverage rate \(r\) no longer cancels. The fee is a fixed sum while the return scales with \(r\), so the amount an author writes now depends on how congested the queue is, and that congestion depends on how much everyone writes. The equilibrium is the point where the two agree. It is unique: a higher \(r\) makes the fee smaller relative to the return, which makes authors write more, which lengthens the queue and lowers \(r\) again. Second, they pay only while the exemption is worth buying. At a given substance choice, paying the fee beats risking desk rejection only when \(P < \mu\, r\,b\,\Lambda_H\). Call this the cap, \(\overline{P} = \mu\, r\,b\,\Lambda_H\): the highest fee the high-quality authors will pay. If the detector rarely catches anyone, nobody will pay to skip it. For those who do pay, the fee acts like an extra drafting cost. It pushes toward fewer, more substantial manuscripts, the opposite of what cheap drafting did in the first-order condition above.
Which fee? Think of the review budget as a road with a fixed number of lanes. Every submission takes a review slot, and when the queue is congested (\(r < 1\)) one more submission pushes some other manuscript out unread. If the pushed-out manuscript was good, the journal loses a good paper, and the author who did the pushing pays nothing for that loss. Economists call this a congestion externality, and the standard remedy is a congestion charge: we bill each author for the damage their submission does to everyone else’s. In this model the charge is easy to write down. In the congested case the throughput is \(Q = M\,A_H/N\), where \(A_H\) is the number of accepted high-quality manuscripts per period and \(N\) is the length of the queue. One more high-quality manuscript adds \(\Lambda_H\) to \(A_H\) and one slot to \(N\), so it changes \(Q\) by \(r\,\Lambda_H - Q/N\). The author counts only the first term, the paper they might get accepted. The second term accounts for the good papers they displace, which costs fall on the other high-quality authors. A fee equal to the second term, converted to money at \(b\) per paper, makes the author’s calculation match the journal’s, and that is the optimal fee.
where \(\sigma_H = \lambda n_H / N\) is the high-quality share of the queue. In words: the return to a high-quality submission, discounted by the share of the queue that is high-quality. The fee must also sit above the floor \(\underline{P}\), so that slop authors don’t pay it, and below the cap \(\overline{P}\), so that high-quality authors do. That gives
\[
P_D \;=\; \min\!\Big\{\max\{P^\circ,\; \underline{P}\},\; \overline{P}\Big\}.
\] Everything on the right depends on the equilibrium, and the equilibrium depends on the fee, so we solve this by numerically iterating to convergence. Figure 6 checks the answer against a brute-force sweep of \(P\). The peak of \(Q\) sits on \(P_D\) at every \((a, \mu)\) I have tried, except where the congestion charge would exceed the cap, in which case the peak sits on the cap. At my parameters, the congestion charge is always above the floor, so the floor never matters. When the queue is uncongested, nobody displaces anybody, the charge is zero, and \(P_D\) drops to the floor, which at \(\mu = 1\) is \(\mu^*\,b\,\Lambda_L\), a fraction \(\mu^*\) of the \(P^*\) of Policy C, because a slop author who cannot slip past the detector still has hand-drafting to fall back on.
Code
def respond_D(r, a, mu, fee, p=P):"""Best responses at coverage r when a fee buys exemption from the classifier. Each manuscript takes one of three routes: hand-draft free, AI-draft and risk the desk, or AI-draft and pay. Returns queue load and accepted count per type.""" delta = p["f"] / a ret = r * p["b"] * accept_H(T_GRID, p) routes = [(ret / (T_GRID + p["f"]), p["f"], 1.0, "hand"), ((1- mu) * ret / (T_GRID + delta), delta, 1- mu, "risk"), ((ret - fee) / (T_GRID + delta), delta, 1.0, "pay")] obj, dH, wH, H_mode =max(routes, key=lambda o: o[0].max()) tH = T_GRID[int(np.argmax(obj))] loadH = p["lam"] * p["T"] / (tH + dH) * wH piL = r * p["b"] * accept_L(p) routes = [(piL / (p["t_min"] + p["f"]), p["f"], 1.0, "hand"), ((1- mu) * piL / (p["t_min"] + delta), delta, 1- mu, "risk"), ((piL - fee) / (p["t_min"] + delta), delta, 1.0, "pay")] _, dL, wL, L_mode =max(routes, key=lambda o: o[0]) loadL = (1- p["lam"]) * p["T"] / (p["t_min"] + dL) * wLreturndict(loadH=loadH, loadL=loadL, aH=loadH * accept_H(tH, p), aL=loadL * accept_L(p), tH=tH, H_mode=H_mode, L_mode=L_mode)def solve_D(a, mu, fee, p=P, iters=60):"""Exemption fee: coverage no longer cancels from anyone's problem, so bisect for the fixed point in r. At a jump in either type's response that type mixes, so blend the two sides of the jump in the proportion that fills the queue.""" M = p["R"] / p["h"] lo, hi =1e-6, 1.0for _ inrange(iters): mid =0.5* (lo + hi) q = respond_D(mid, a, mu, fee, p)ifmin(1.0, M / (q["loadH"] + q["loadL"])) > mid: lo = midelse: hi = mid r =0.5* (lo + hi) qlo, qhi = respond_D(lo, a, mu, fee, p), respond_D(hi, a, mu, fee, p) Nlo, Nhi = qlo["loadH"] + qlo["loadL"], qhi["loadH"] + qhi["loadL"] congested = r <1-1e-6 alpha = np.clip((M / r - Nhi) / (Nlo - Nhi), 0, 1) if congested andabs(Nlo - Nhi) >1e-9else1.0 mix =lambda k: alpha * qlo[k] + (1- alpha) * qhi[k] aH, aL, N = mix("aH"), mix("aL"), mix("loadH") + mix("loadL") sH = p["theta"] * (alpha * qlo["aH"] * np.sqrt(qlo["tH"])+ (1- alpha) * qhi["aH"] * np.sqrt(qhi["tH"])) / aH W = aH * sH / (aH + aL) if aH + aL >0else0.0returndict(r=r, W=W, thruH=r * aH, tH=mix("tH"), sH=sH, N=N, H_pays=qlo["H_mode"] =="pay", floor=min(mu, mu_star(a, p)) * r * p["b"] * accept_L(p), charge=p["b"] * r * aH / N if congested else0.0)def fee_D(a, mu, p=P, eps=1e-3, iters=200):"""Q-optimal fee: the fixed point of P = max(b Q / N, floor + eps), capped at the largest fee the high-quality authors will pay rather than risk the desk.""" fee = p["b"] * accept_L(p)for _ inrange(iters): q = solve_D(a, mu, fee, p) new =max(q["charge"], q["floor"] + eps)ifabs(new - fee) <1e-7:break fee =0.5* (fee + new)ifnot solve_D(a, mu, fee, p)["H_pays"]: lo, hi =0.0, feefor _ inrange(40): mid =0.5* (lo + hi) lo, hi = (mid, hi) if solve_D(a, mu, mid, p)["H_pays"] else (lo, mid) fee = loreturn fee
Code
from matplotlib.lines import Line2Dfees = np.linspace(0.0, 0.5, 251)fig, (axq, axw) = plt.subplots(1, 2, figsize=(8, 3.2))for mu, col in ((0.6, "#2d6e8e"), (0.95, "#3a7d44")): sweep = [solve_D(100, mu, f) for f in fees] axq.plot(fees, [q["thruH"] for q in sweep], color=col, label=f"$\\mu = {mu}$") axw.plot(fees, [q["W"] for q in sweep], color=col) fd = fee_D(100, mu)for ax_ in (axq, axw): ax_.axvline(fd, color=col, lw=0.8) ax_.axvline(solve_D(100, mu, fd)["floor"], color=col, ls=":", lw=0.8)axq.set_ylabel("good papers per author-period $Q$")axw.set_ylabel("scientillas per article $W$")for ax_ in (axq, axw): ax_.set_xlabel("fee $P$")handles, labels = axq.get_legend_handles_labels()handles += [Line2D([], [], color="#555555", lw=0.8), Line2D([], [], color="#555555", ls=":", lw=0.8)]labels += ["$P_D$", r"floor $\underline{P}$"]axq.legend(handles, labels, fontsize=8)fig.tight_layout()plt.show()
Figure 6: The exemption fee at \(a = 100\): throughput \(Q\) (left) and scientillas per article \(W\) (right) against the fee \(P\), at two true-positive rates. Solid verticals mark the formula \(P_D\); dotted verticals mark the floor \(\underline{P}\), below which slop authors pay the fee too. The cliff on the right of each curve is where high-quality authors stop paying and take their chances with the classifier.
Code
mus_d = np.linspace(0.05, 0.98, 60)fig, axp = plt.subplots(figsize=(6, 3.4))for a, col inzip((2, 5, 15, 100), ("#b5541c", "#7f4c94", "#2d6e8e", "#3a7d44")): fd = [fee_D(a, mu) for mu in mus_d] axp.plot(mus_d, fd, color=col, label=f"$a = {a}$") axp.plot(mus_d, [solve_D(a, mu, f)["floor"] for mu, f inzip(mus_d, fd)], color=col, ls="--", lw=0.8)handles, labels = axp.get_legend_handles_labels()handles += [Line2D([], [], color="#555555", lw=1.2), Line2D([], [], color="#555555", ls="--", lw=0.8)]labels += ["optimal fee $P_D$", r"floor $\underline{P}$"]axp.set_xlabel(r"classifier true-positive rate $\mu$")axp.set_ylabel("fee")axp.legend(handles, labels, fontsize=8, ncol=2)fig.tight_layout()plt.show()
Figure 7: The optimal exemption fee \(P_D\) against classifier true-positive rate \(\mu\), at several capability levels, with the floor \(\underline{P}\) dashed. The fee rises with the true-positive rate while AI-drafted slop manuscripts remain in the queue, is flat in it once slop authors have gone back to hand-drafting above \(\mu^*\), and at low true-positive rates is capped by what skipping the classifier is worth to a high-quality author.
The classifier’s true-positive rate doesn’t appear in the congestion charge. In Policy D, the classifier does one thing: it gives AI-drafting authors a reason to pay by desk-rejecting a fraction \(\mu\) of the manuscripts that do not. The fee amount depends on how congested the review queue is, not on the classifier. The true-positive rate does enter the floor and the cap, and both rise with \(\mu\), because skipping the classifier is worth more to an author when the classifier is better. So, a better detector widens the range of fees that high-quality authors will pay and slop authors will not, while a worse one narrows it. At a low enough true-positive rate, the cap falls below the congestion charge. The classifier then catches so few manuscripts that skipping it is worth little, and the journal can charge only what high-quality authors will still pay, which is less than the congestion charge. Above \(\mu^*\), Figure 7 is flat in \(\mu\), because slop authors have gone back to drafting by hand, so the classifier flags nothing and the length of the queue no longer depends on it. Unlike \(P^*\), this fee changes as \(a\) grows, because \(Q\) and \(N\) do.
The two welfare metrics disagree here. \(W\) prefers the largest fee the high-quality authors will pay, right up to the cap. That is because \(W\) is blind to publication volume, so long ast the average quality is high. Both metrics agree that the exemption fee beats the ban at every \((a, \mu)\) I computed, so the ranking of policies is unchanged; they disagree rather on how high to set the fee. Near \(a = 1\), the fee pushes high-quality authors back to hand-drafting, because AI saves them almost nothing there, and Policy D degenerates into Policy B.
Code
fig, (ax, axt) = plt.subplots(1, 2, figsize=(9.5, 3.4))eqD = {mu: [solve_D(a, mu, fee_D(a, mu)) for a in a_grid] for mu in (0.6, 0.95)}for ax_, key in ((ax, "W"), (axt, "thruH")): ax_.plot(a_grid, [q[key] for q in eqA], color="#b5541c", label="A: free-for-all")for mu, ls in ((0.6, "--"), (0.95, "-")): ax_.plot(a_grid, [solve_B(a, mu)[key] for a in a_grid], ls, color="#2d6e8e", label=f"B: classifier, $\\mu={mu}$") ax_.plot(a_grid, [q[key] for q in eqD[mu]], ls, color="#7f4c94", label=f"D: exemption fee, $\\mu={mu}$") ax_.plot(a_grid, [q[key] for q in eqC], color="#3a7d44", label="C: submission fee") ax_.set_xscale("log") ax_.set_xlabel("AI capability $a$")ax.set_ylabel("scientillas per article $W$")axt.set_ylabel("good papers per author-period $Q$")ax.legend(fontsize=7)fig.tight_layout()plt.show()
Figure 8: The four policies as AI capability grows. Classifier and exemption fee at true-positive rate \(0.6\) and \(0.95\); submission fee at \(P^*\); exemption fee at \(P_D\). Left: scientillas per published article \(W\). Right: good papers published per author-period \(Q\). The exemption fee sits between the ban and the submission fee at every capability level \(a\).
The exemption fee lands between the ban and the submission fee on both welfare metrics, at every capability level in Figure 8. Above \(\mu^*\) the ban restores the hand-drafting equilibrium, \(Q \approx 0.22\) good papers per author-period at my parameters, regardless of \(a\). Under the exemption fee at the same true-positive rate the journal publishes \(0.39\) good papers per author-period at \(a = 100\) (39 a period per hundred authors, against 22 under the ban), because high-quality authors keep the time AI saves them and the congestion charge stops them from spending all of it on extra manuscripts. Below \(\mu^*\), at \(\mu = 0.6\), the exemption fee roughly doubles the ban’s throughput, but both are poor, because the \(40\%\) of AI-drafted slop manuscripts that the classifier misses still fill the queue. At high true-positive rates the journal also publishes more papers in total under the exemption fee than under any other policy, \(0.53\) per author-period against \(0.50\) under the submission fee, because hand-drafted slop is still in the queue and some of it still passes review; it loses to the submission fee on quality, not on volume. What separates the exemption fee from the submission fee is hand-drafted slop. Under Policy C slop authors pay the fee whether or not they used AI, so they stop submitting. Under Policy D a slop author who drafts by hand pays nothing and is never flagged, so their manuscripts stay in the queue. Because of them, the hand-drafting equilibrium’s count of slop manuscripts in the queue is the lowest Policy D can reach, just as it was the lowest Policy B could reach; the slop share of the published journal still falls below that equilibrium’s, because the high-quality authors submit more. At \(\mu = 1\) an unpaid AI-drafted manuscript is certain to be desk-rejected, so every AI drafter pays, and Policy D is Policy C levied on AI drafters alone. The slop authors who draft by hand are the whole difference.
For the Alignment Journal the exemption fee faces the same objection as the submission fee, that we would be charging money from a cash-constrained author base. Only authors who choose to pay are charged. Paying is also a disclosure: a hand-drafted manuscript is never flagged, so the only reason to buy the exemption is that the manuscript was AI-drafted. Most journals now ask authors to declare AI use and but can at best partially check the claim; here it arrives with a fee attached, so it is truthful by construction. A false positive still costs a hand-drafting author their manuscript, as under the ban. And the revenue arrives in proportion to the number of AI-drafted manuscripts the journal has to review, so a journal that pays its reviewers could route it straight back into \(R\).
6 Future work
On this analysis, and regarded purely as an instrument for controlling slop rates, the classifier-based ban is a clumsy approximation to a congestion charge. Policy D makes the charge explicit, and the formula for it says what we are charging for: constrained review slots. The ban is imperfect: as its true-positive rate improves, it starts hurting the compliant at \(\mu_H\) and only starts working as \(\mu\) increases past \(\mu^*\). Whether the Alignment Journal, or any journal, has a reasonable economic basis to adopt it depends on some concrete parameters that I have only guessed here:
The true-positive rate a detector can sustain against motivated rewriting, compared against a \(\mu^*\) we could estimate from the time costs of producing slop; and
how many good AI-drafted papers we would tolerate losing in the range of true-positive rates between the two thresholds.
6.1 Diluted review
Hereto I have assumed that the journal holds the review depth \(h\) fixed and rations the coverage proportion \(r\) by skipping some papers. The failure in such a setting is not too badly behaved, in that the equilibrium is unique, welfare falls continuously in \(a\), and a marginal improvement — a little more budget, a slightly better classifier — buys a proportionate marginal repair.
We could contrariwise imagine the journal absorbs a growing volume of submissions by reviewing all of them less carefully: coverage is fixed \(r = 1\), and each manuscript gets \(e = R/N\) reviewer-hours instead of the fixed \(h\). That puts congestion “inside” the incentives: the return to substance is now scaled by the per-manuscript attention \(e\), so as review thins, even high-quality authors would stop investing in the substance of manuscripts, which frees more time for volume, which thins review further. That feedback loop produces weird, pathological behaviour, depending on the proportion of authors of type \(H\) and \(L\).
Code
T2 = np.geomspace(P["t_min"], 25.0, 600)def tH_dilute(e, delta, p=P):"""H's optimal substance time when acceptance depends on attention e.""" obj = logistic(e * (p["theta"] * np.sqrt(T2) - p["sbar"])) / (T2 + delta)return T2[int(np.argmax(obj))]def dilution_equilibria(a, p, n=600):"""All fixed points of e * N(e) = R in the dilution variant.""" delta = p["f"] / a nL = p["T"] / (p["t_min"] + delta)def load(e): tH = tH_dilute(e, delta, p)return p["lam"] * p["T"] / (tH + delta) + (1- p["lam"]) * nL es = np.geomspace(0.1, 50.0, n) g = np.array([e * load(e) - p["R"] for e in es]) roots = []for i inrange(n -1):if (g[i] <0) != (g[i +1] <0): lo, hi = es[i], es[i +1]for _ inrange(40): mid = np.sqrt(lo * hi)if ((mid * load(mid) - p["R"]) <0) == (g[i] <0): lo = midelse: hi = mid roots.append(np.sqrt(lo * hi))return rootsfold_a = np.concatenate([np.linspace(1.5, 3.1, 20), np.linspace(3.1, 3.7, 45), np.linspace(3.7, 8.0, 20)])fig, axes = plt.subplots(1, 2, figsize=(9.5, 3.4), sharey=True)for ax, lam inzip(axes, (0.25, 0.75)): p2 =dict(P, lam=lam)for a in fold_a:for e in dilution_equilibria(a, p2): ax.plot(a, e, "o", ms=3, color="#7f4c94") ax.set_xlabel("AI capability $a$") ax.set_title(f"$\\lambda = {lam}$")axes[0].set_ylabel("equilibrium attention $e$")fig.tight_layout()plt.show()
Figure 9: Equilibrium attention \(e\) at each capability \(a\), in the variant where the review depth \(h\) is spread across the whole queue, everything else as in the baseline. Left: at our baseline mix \(\lambda = 0.25\) the equilibrium is unique everywhere, so the decline is a slope. Right: at \(\lambda = 0.75\) a narrow range of \(a\) has three equilibria — a high-attention branch, a collapsed branch, and an unstable one between — so the decline is a cliff.
Figure 9 illustrates one case of each. At our baseline author mix, \(\lambda = 0.25\), the dilution variant also declines smoothly: one equilibrium at every \(a\). At \(\lambda = 0.75\) there is a narrow range of capability values, around \(a \approx 3.4\), in which three equilibria exist at once; the middle one is unstable, so in practice the journal is in one of two self-consistent states: careful review with restrained submission, or cursory review with a flood. Which state it ends up in depends on its history. Suppose the journal is in the careful state and \(a\) rises through this range. The journal stays careful, changing only a little — until, at the top of the range, the careful state stops being self-sustaining at all, and the journal falls to the flooded state in one step. The fall does not reverse: lower \(a\) back into the range and the journal stays flooded, because the flooded state is self-sustaining there too. Nobody can lower \(a\), but the model only sees \(a\) through the drafting cost \(\delta = f/a\), and a journal can raise that: a fee adds to the cost of every manuscript, and a ban above \(\mu^*\) resets it to \(f\). Recovery means raising the effective cost of a manuscript well past the level at which the fall began, to where the flooded state itself stops existing, or else expanding the review budget \(R\) enough to move the fold. Bartolucci and Vivo (2026) works this margin out properly in a queueing model: under load, reviewers rationally raise the risk threshold for checking AI output, cutting scrutiny when it matters most.
6.2 More capable AI
Slop’s citoms \(s_L\) are fixed in this model; the realistic case has them rising with \(a\), and as \(s_L \to \bar{s}\) the range of separating fees closes — at which point the journal’s problem stops being mechanism design and becomes epistemology. More broadly, everything here describes the world as it stands: the premise that AI can write a paper but not do the science behind it is a statement about current capability, and the measured trend is that the boundary moves; none of these results survive the regime where it fails. The high-quality authors’ half of the wedge is assumed away artificially: citoms and scientillas coincide for them by assumption, so nobody can oversell. If AI drafting also made real but modest work read as better than it is, the high-quality authors would gain an inflation margin of their own. 🚧TODO🚧
6.3 Harmful slop
Accepted slop is harmless here; if it instead poisons training corpora (Shumailov et al. 2023) or locks in error (Qiu et al. 2025), its value is negative and everything above understates the case for keeping it out. 🚧TODO🚧
6.4 Many journals
There is, in this model, only one journal; with many journals, slop authors send their manuscripts to whichever journal is least strict and there is a race to the bottom, which is a different and worse game. 🚧TODO🚧
6.5 False positives
The classifier is assumed to have a false-positive rate of zero, which is unlikely in practice.
🚧TODO🚧
6.6 Reputations
If publishing slop were reputationally harmful, the return to slop would change and the amount of spam might also be moderated.
🚧TODO🚧
7 Ballpark real-world estimates
The model prices everything in units of \(b\), the value of one accepted paper to its author, so a fee in dollars needs a guess at \(b\) in dollars. My best anchor is cost: a typical ICLR paper has historically taken about four researcher-months, and a careerist keeps spending that only if an acceptance is worth at least as much to them, so four researcher-months is a lower bound on \(b\). At USD 5,000 to USD 15,000 per researcher-month, from a PhD stipend to a fully loaded academic salary, that puts \(b\) at USD 20,000 to USD 60,000. I will carry both ends through.
Under Policy C the fee at which slop authors stop submitting is \(P^* = b\,\Lambda_L\), which at my parameters is about an eighth of \(b\): USD 2,500 to USD 7,500 per submission. Whatever it is, that does not change as the AI improves.
Under Policy D with a Pangram-grade classifier, taking the true-positive rate from credible independent tests (Russell, Karpinska, and Iyyer 2025) as \(\mu = 0.993\) (i.e. \(99.3\%\)) with no false positives, that rate clears the large-\(a\) limit, i.e. \(\mu^*(a)\simeq 0.91<\mu^*\) for every \(a\), so slop authors draft by hand, take their chances and never pay the fee. The fee is then a pure congestion charge. It rises slowly with \(a\), from \(0.16\,b\) at \(a = 2\) to \(0.20\,b\) at \(a = 100\). Call it a fifth of \(b\): USD 4,000 to USD 12,000 per AI-drafted submission. The floor is \(0.05\)–\(0.09\,b\), well below that, so slop authors do not pay at any fee in this range. For comparison, at the same true-positive rate the ban restores the hand-drafting equilibrium, \(0.22\) good papers per author-period, while under the exemption fee the journal publishes \(0.39\) good papers per author-period at \(a = 100\): 22 against 39 a period for every hundred authors.
The ban has a price too, paid in time rather than money. Above \(\mu^*\) every high-quality author drafts by hand, so each of their manuscripts costs \(f - \delta\) more time than it would with AI, about \(0.5\) time units at \(a = 100\). On the same anchor, a paper takes about \(0.96\) time units and four researcher-months, so the ban costs a good author about two researcher-months per manuscript: roughly half of \(b\), or USD 10,000 to USD 30,000. That is between two and three times the exemption fee, and the journal does not even collect it. A second way to price the ban gives the same order: the exemption fee at which a high-quality author would be as well off as under the ban is \(0.44\,b\) at \(a = 100\), falling to \(0.19\,b\) at \(a = 2\), where AI saves little drafting time and the ban costs little.
The exemption fee face value is higher than the submission fee. Under Policy C the fee only has to keep slop authors out, and once they are gone the queue is uncongested, so there is no congestion to charge for. Under Policy D the slop authors who draft by hand stay, the queue stays congested, and the fee is a congestion charge on the high-quality authors.
Both fees are large next to what journals that do charge actually ask; economics journals sit around USD 100–300 per submission, I believe. Either the going rate is low, or \(b\) is smaller than I have guessed for most authors, or my \(\Lambda_L\) is too high. A journal whose full review lets through fewer than \(12\%\) of slop manuscripts needs a proportionately smaller fee, since \(P^*\) scales with \(\Lambda_L\).
We could be cute The model treats slop as zero scientillas rather than hypothesizing slopons, the antiparticle of the scientilla, which annihilates one on contact over the course of a reader’s afternoon; see Harmful slop.↩︎
The classifier reads style, not substance, and the two are coming apart: in the neighbouring market for cover letters, an AI writing tool cut the correlation between text quality and callbacks by half (Cui, Dias, and Ye 2025).↩︎