Arvind Narayanan · Defending institutions against AI slop

Taxonomy of Institutional Responses to AI-Generated Floods

How institutions are responding when AI drops the cost of an interaction to near-zero, threatening to overwhelm human-scale review processes. Synthesized from research across 50 interaction types.

How to read this file

Exhibits within each category are ranked, best first, on three criteria combined:

  1. Importance — how much institutional weight the example carries (a binding policy at a major venue beats a vendor blog).
  2. Relevance — how directly it demonstrates this response category rather than adjacent noise.
  3. Persuasiveness — verifiable numbers, a named institution, a primary source, and a mechanism a skeptical reader can check.

Tiers:

⚠️ = claim not fully verified against a primary source, or sourced from an interested party. See Source quality at the end.


The 13 Response Categories

1. AI Detection Tools (fighting AI with AI)

Still the most common response, and the one most often described wrongly. Its reliability is sharply bimodal: there are two populations of detector, and treating them as one class is what produces most bad claims in this area — in either direction.

Frontier trained classifiers Legacy / commodity detectors
Examples Pangram, EditLens GPTZero, ZeroGPT, Turnitin's AI-writing feature, dozens of SEO-market tools
Reported FPR ~1 in 10,000 overall; 1 in 100,000 on held-out arXiv papers ⚠️ (vendor self-report, third-party-validated) Turnitin: <1% document-level, <4% sentence-level (vendor self-report)
Behavior under domain shift Not independently characterized Degrades substantially (Pudasaini et al. 2026)
Institutional use Population-scale audits; desk rejections at NeurIPS Being switched off across higher ed

Ranked exhibits:

On false positives and equity. The claim that "false positives disproportionately harm non-native English speakers" is widely repeated as a present-tense fact. It was established on 2023 perplexity detectors, and three things need stating separately:

  1. The original result is real but narrow. Liang et al. (2023) found ~61% FPR on TOEFL essays vs. 1–4% on native-speaker essays across seven perplexity-based detectors. The mechanism was linguistic simplicity, not nationality: enriching the same TOEFL essays to sound more native-like cut the FPR from 61.2% to 11.8%. Any writer with plain syntax is exposed — L2 writers just are, disproportionately.
  2. Frontier classifiers do not obviously reproduce it, and no one has published the equivalent test on them. That is a genuine gap, not a clean bill of health. The absence of a Liang-style audit of Pangram is the single most useful missing study in this area.
  3. The harm has migrated from error rate to base rate. At population scale a correct detector still produces many wrongful accusations, because the denominator is enormous and the accused have no way to prove a negative. NeurIPS understood this precisely — hence the appeal route requiring version history with pre-AI and post-AI checkpoints. Note what that demands: an audit trail most researchers do not keep, which is easiest to produce for people working in well-resourced environments with modern tooling. The equity problem did not disappear; it moved from "whose prose looks robotic" to "who can document their process." That is a better problem, and still a real one.

2. Identity Verification & Authentication

Proving humanness rather than detecting AI text. Durable in principle; adoption is the constraint.


3. Volume Caps & Rate Limiting

Blunt, fast, and effective against quantitative flooding only.

Structural limit: rate limits do nothing about qualitative flooding. Schmitz et al. find complexity increases in 90% of their 84 government cases versus volume increases in 60% — including a single German social-court letter running over 4,000 pages. A cap of six NIH applications does not stop each of the six from being longer and denser than any human would have written unaided. Most of this flood is not more submissions; it is heavier ones.


4. Disclosure & Transparency Requirements

Cheap to enact, and the category whose weakness the 2026 evidence exposes most clearly.

Assessment: disclosure is not a response to flooding so much as a precondition for other responses, and it is only as strong as the verification behind it. Unenforced, it selects against honest declarers.


5. Outright Bans & Shutdowns

The nuclear option. Increasingly chosen not by cranks but by well-run projects that did the arithmetic.

Note the asymmetry: every one of these closes a channel that was open because openness was cheap. The bug bounty, the unsolicited submission, the drive-by PR, the anonymous edit — these were gifts from a world where producing a plausible artifact was expensive. Bans are institutions revoking the gift.


6. Process Redesign (shifting to harder-to-fake interactions)

The most durable category, and the best-evidenced.


7. Certification & Provenance

Marks of human origin. Still the weakest-adopted durable idea — but see the NeurIPS entry in §6, which is provenance arriving through the back door.

Why this stays low: certification asks the honest to bear a cost the dishonest can decline. It works only where the certificate gates something valuable, which is why the NeurIPS audit trail — where the "certificate" is the price of appealing a rejection — is more likely to stick than a voluntary badge.


8. Legislation & Regulation

Slowest, most authoritative, and currently going backwards in the US.


9. Financial Penalties & Sanctions

Reactive, but unusually well measured.


10. Industry Collaboration & Shared Infrastructure

Pooling detection and verification costs. Durable, underrated, quietly the model that scales.


11. Co-optation & Platform Integration

If you can't beat them, absorb them. Solves volume; changes what the interaction means.


12. Education & Behavioral Adaptation

Training humans to cope. Cheapest, weakest, but occasionally the only available move.


13. Reduce the Expected Benefit

Make the flood not worth generating, by shrinking what a successful submission wins. Drawn from Schmitz, Hammond & Chan, whose four-class response framework treats it as a first-order strategy alongside adding friction.

Why this category matters more than its exhibit count suggests: it is the response with the worst distributive consequences and the least visibility. A cap is legible and contestable. A quietly narrowed eligibility rule is neither. If the project has a policy-relevant warning to give, it is probably here.


Summary: Which Responses Are Most Durable?

Detection cannot be ranked as a single thing; it splits in two, and the split does most of the work here.

Most durable:

Detection, split:

Moderately durable:

Least durable:

Emerging and unproven:


Key Cross-Cutting Findings

  1. No institution has fully solved it. Every response remains partial, and most remain reactive.

  2. The detection arms race is not one race. The familiar claim — detection is always one generation behind — held for perplexity-based methods. Trained classifiers with published negative controls are a different instrument, and in 2026 they carried institutional weight for the first time. The honest formulation: detection is now good enough to measure a population and too contestable to convict an individual without corroboration. NeurIPS is the proof of both halves — it acted on the measurement, and it still built an appeals process.

  3. Qualitative flooding is the bigger half, and most responses only touch the quantitative half. Complexity increases appear in 90% of Schmitz et al.'s cases, volume increases in 60%. Rate limits, caps, and CAPTCHAs address volume. A 4,000-page social-court letter defeats all of them.

  4. The equity cost has moved, not vanished. From "whose prose reads as machine-like" (Liang et al. 2023, perplexity detectors) to "who can produce a documented audit trail" (NeurIPS 2026) and "who can afford the reintroduced fee or the in-person appearance" (Schmitz et al.). Each successive response is fairer than the last on its own terms and still sorts by resources. This is the finding most worth carrying into the paper.

  5. Process redesign is underused because it is expensive, not because it is unknown. Institutions are not defaulting to detection out of ignorance; they are defaulting to it because oral exams cost faculty hours per student and service redesign takes years, while a classifier costs a licence fee.

  6. Some floods are beneficial, and the beneficial ones are hardest to filter. AI-assisted CFPB complaints, FOIA requests, benefits appeals, and tax challenges democratize access. Schmitz et al. put this most sharply: agents genuinely reduce administrative burden and unlock legitimately suppressed demand, so the problem is capacity, not demand — and friction-based responses "lock this potential further out of reach." Note the tension with §13: reducing the expected benefit is precisely a decision to suppress demand rather than build capacity.

  7. Physical presence and provenance are the two things AI cannot cheaply fake. In-person interviews, oral defences, blue books — and, increasingly, version history. Both are bets on evidence that exists outside the artifact.

  8. The "both sides use AI" equilibrium is now measurable. FOIA (~20% of agencies process with AI), peer review (AAAI's 22,977 AI-assisted reviews; NeurIPS's randomized trial), RFPs, hiring, planning objections. In several domains this looks stable but is not a return to the status quo: both sides now spend more to reach the same decision.

  9. The successful responses share a structuremeasure first, act second, and publish the calibration. NeurIPS validated Pangram against a pre-ChatGPT control before rejecting anyone. curl counted its confirmation rate before killing the bounty. Wikipedia specified observable signatures before authorizing speedy deletion. The failures share the opposite structure: adopt a tool, act on its output, discover the error rate from the complaints.


Source quality

Read against primary sources: the NeurIPS desk rejections and methodology; the Pangram ICLR analysis; the curl bug-bounty closure; the NIH cap and AI policy; Curtin's detection disablement; FOIA volumes and agency AI use (CJR/Tow); Pudasaini et al.; Liang et al.; Spotify removals; the Verisk insurance-fraud survey; and Wikipedia's G15 policy text.

⚠️ Not verified — do not treat as load-bearing:

Known gap: no published Liang-style demographic false-positive audit of a frontier classifier. Until one exists, claims that modern detection is equitable are unsupported in the same way that claims it is inequitable are out of date.

A note on the supporting files. The five responses_*.md files contain useful per-type detail but carry no citations at all — roughly 1,500 lines with zero URLs between them. Anything promoted from them into this document is flagged above. They should be re-sourced before the project is published; until then, this file is the citable layer and they are working notes.