How institutions are responding when AI drops the cost of an interaction to near-zero, threatening to overwhelm human-scale review processes. Synthesized from research across 50 interaction types.
How to read this file
Exhibits within each category are ranked, best first, on three criteria combined:
- Importance — how much institutional weight the example carries (a binding policy at a major venue beats a vendor blog).
- Relevance — how directly it demonstrates this response category rather than adjacent noise.
- Persuasiveness — verifiable numbers, a named institution, a primary source, and a mechanism a skeptical reader can check.
Tiers:
- ★★★ Lead exhibit — Primary source, named institution, checkable number.
- ★★ Supporting — solid corroboration; use to show the pattern is not a one-off.
- ★ Background — real but thin, dated, secondary, or interested-party.
⚠️ = claim not fully verified against a primary source, or sourced from an interested party. See Source quality at the end.
The 13 Response Categories
1. AI Detection Tools (fighting AI with AI)
Still the most common response, and the one most often described wrongly. Its reliability is sharply bimodal: there are two populations of detector, and treating them as one class is what produces most bad claims in this area — in either direction.
| Frontier trained classifiers | Legacy / commodity detectors | |
|---|---|---|
| Examples | Pangram, EditLens | GPTZero, ZeroGPT, Turnitin's AI-writing feature, dozens of SEO-market tools |
| Reported FPR | ~1 in 10,000 overall; 1 in 100,000 on held-out arXiv papers ⚠️ (vendor self-report, third-party-validated) | Turnitin: <1% document-level, <4% sentence-level (vendor self-report) |
| Behavior under domain shift | Not independently characterized | Degrades substantially (Pudasaini et al. 2026) |
| Institutional use | Population-scale audits; desk rejections at NeurIPS | Being switched off across higher ed |
Ranked exhibits:
- ★★★ NeurIPS 2026 Position Paper Track desk-rejects 178 papers (18.4%) using Pangram. NeurIPS blog, 2 June 2026. Organizers ran every submission through Pangram v3.3.2, found 28.2% scored 100% under default windowing, then did the work almost nobody does: they validated the instrument before acting on it. Negative control on FAccT 2022 papers (pre-ChatGPT, stylistically comparable) returned 0.0% at every threshold; FAccT 2025 returned 1.0% at ≥50% and 0.0% at 100%. They re-ran with smaller text windows to avoid over-claiming (dropping ≥90% scores from 42.7% to 12.7%), tested 12 categories of AI use against their own policy (all permissible uses — proofreading, light copyediting, translation — went unflagged; all clearly impermissible ones were flagged), and found partial AI completions of ≤20% were never flagged. Final action: 178 desk rejections without appeal, 123 more conditional on producing a documented pre-AI/post-AI audit trail. This is the strongest available evidence that detection can be made decision-grade — and a template for what an institution has to do first.
- ★★★ Pangram audit of ICLR 2026: 21% of 76,139 peer reviews fully AI-generated. Pangram, Nov 2025; covered by Nature. Over half of reviews showed some AI involvement. Negative control on ICLR 2022 reviews gave error rates of 1 in 1,000 (light edit vs. human) to 1 in 10,000 (heavy edit vs. human), with no confusions between fully-AI and fully-human. Population-scale measurement of an institution's integrity, done by an outside party, for a bounty — a genuinely new capability. ⚠️ Vendor-run study of its own product; the negative controls are the reason to take it seriously anyway.
- ★★ Universities are switching detection off, not up. Curtin University disabled Turnitin's AI-writing detection across all campuses from 1 January 2026, citing reliability, equity, and a preference for assessment redesign over surveillance; text-matching stays on. Vanderbilt disabled it in 2023. ⚠️ A widely repeated "more than 50 institutions" figure traces only to detector-marketing sites — treat the trend as real and the count as unverified.
- ★★ Benchmark accuracy is not authorship detection. Pudasaini et al., "Why AI-Generated Text Detection Fails" (arXiv:2603.23146, Apr 2026). Linguistic-feature models hit F1 0.9734 on PAN-CLEF 2025 and 0.8025 on COLING 2025, then degrade sharply under cross-domain and cross-generator shift. SHAP analysis shows the most influential features differ markedly between datasets — detectors are keying on corpus artifacts, not machine authorship. The authors' framing is the useful one: a detector is valid only if the evidence it uses stays stable under domain and generator shift. Applies to the feature-based family they test; it is the right prior for any detector that has not published negative controls.
- ★★ Spotify removed 75 million spam tracks in twelve months and is shipping an AI music spam filter in autumn 2026 (ABC News, 16 July 2026). Detection at a scale where human review was never an option.
- ★★ Amazon blocked 250M+ suspected fake reviews in 2025 using ML, NLP, and graph neural networks; Google's Gemini-powered review scanning blocks 85%+ of fake reviews before publication, with review deletions up 600% in H1 2025 — and a documented false-positive problem, legitimate five-star reviews disproportionately removed. ⚠️ Both company self-reports. The Google case is the useful one: it is the clearest instance of detection at scale visibly harming legitimate participants.
- ★ Publisher-side screening: Wiley Papermill Detection (6 mechanisms, ~10K manuscripts/month); STM Integrity Hub (40 publishers, 125K papers/month). Real infrastructure, but see §10 — its value is as shared infrastructure more than as detection.
- ★ ICML 2026 prompt-injection trap: hidden instructions in submitted PDFs trigger telltale phrases if a reviewer feeds the paper to an LLM; hundreds of reviewers caught. Ingenious, contested on ethics, and works by exploiting a current model limitation rather than a durable principle — so it belongs in "emerging and unproven," not here.
- ★ Journal-level detection in practice: Clinical Orthopaedics and Related Research flagged 21 of 43 suspicious letters using AI detectors; PRiMER deployed QuillBot, GPTZero and ZeroGPT. Concrete, small-scale, and honest about the tools' error rates — a useful counterpoint to the venue-scale NeurIPS and ICLR numbers.
- ★ Sector adoption figures: 65% of insurers use automated AI fraud detection; 62% of scholarship providers use AI detection (up from 28%); 68% of schools integrated detectors. ⚠️ All from secondary/industry sources, none traced to a primary; useful only as texture.
On false positives and equity. The claim that "false positives disproportionately harm non-native English speakers" is widely repeated as a present-tense fact. It was established on 2023 perplexity detectors, and three things need stating separately:
- The original result is real but narrow. Liang et al. (2023) found ~61% FPR on TOEFL essays vs. 1–4% on native-speaker essays across seven perplexity-based detectors. The mechanism was linguistic simplicity, not nationality: enriching the same TOEFL essays to sound more native-like cut the FPR from 61.2% to 11.8%. Any writer with plain syntax is exposed — L2 writers just are, disproportionately.
- Frontier classifiers do not obviously reproduce it, and no one has published the equivalent test on them. That is a genuine gap, not a clean bill of health. The absence of a Liang-style audit of Pangram is the single most useful missing study in this area.
- The harm has migrated from error rate to base rate. At population scale a correct detector still produces many wrongful accusations, because the denominator is enormous and the accused have no way to prove a negative. NeurIPS understood this precisely — hence the appeal route requiring version history with pre-AI and post-AI checkpoints. Note what that demands: an audit trail most researchers do not keep, which is easiest to produce for people working in well-resourced environments with modern tooling. The equity problem did not disappear; it moved from "whose prose looks robotic" to "who can document their process." That is a better problem, and still a real one.
2. Identity Verification & Authentication
Proving humanness rather than detecting AI text. Durable in principle; adoption is the constraint.
- ★★★ Reddit deploying passkeys, biometrics, World ID, and government IDs for suspected bots (March 2026) — a platform whose entire value is human conversation concluding that content analysis is insufficient.
- ★★ World ID reports 18M+ verified humans and 450M+ verifications, with integrations announced across Zoom, DocuSign, and Shopify. ⚠️ Company figures, crypto-adjacent, unaudited — the category (proof-of-personhood as infrastructure) matters more than the numbers.
- ★★ Email authentication as the working precedent: Google/Yahoo/Microsoft enforcing SPF, DKIM, DMARC; Gmail rejecting non-compliant bulk senders since Nov 2025. The one place where identity verification actually became universal — worth studying as the model.
- ★★ Digital identity as the recommended government fix. Schmitz, Hammond & Chan (2026) name integrating digital identity into the most exposed services as one of three near-term recommendations.
- ★ Dating apps: Tinder FaceCheck, 3D liveness checks, mandatory photo verification at Bumble — verification of the person replacing scrutiny of the photo, now that all three major apps permit AI-generated images.
- ★★ LinkedIn verified 55M users, made verification mandatory for brands, HR teams and executive accounts (late 2025), and reports detecting 99.65% of fake accounts proactively — 97% before any user reports them. It also bans AI-generated profile pictures lacking real-world metadata. ⚠️ Platform self-report. If accurate it is the strongest identity-verification number anywhere in this taxonomy, and worth tracing to LinkedIn's transparency reporting.
- ★ Regulations.gov: CAPTCHA + API rate limits — really §3 wearing an identity costume.
3. Volume Caps & Rate Limiting
Blunt, fast, and effective against quantitative flooding only.
- ★★★ NIH caps PIs at 6 applications per calendar year (NOT-OD-25-132, effective Sept 2025). Still the cleanest single instance of an institution changing its rules because AI broke a cost assumption. The trigger is the persuasive part: some PIs submitted more than 40 applications in a single round. Note the cost — only 1.3% of applicants exceeded six in 2024, so the cap constrains a tiny minority and imposes a ceiling on everyone.
- ★★ GitHub shipped a pull-request kill switch (Feb 2026): disable PRs entirely, or restrict to collaborators. Platform-level acknowledgment that open contribution has an unpriced cost.
- ★ Somerset County (PA) banned anonymous records requests entirely.
- ★ LinkedIn comment-rate limits; email spam-complaint thresholds (<0.3%).
- ★ Regulations.gov CAPTCHA plus API rate limits (50 requests/minute, 500/hour), with GSA able to revoke API access. These failed against CiviClick, because the campaign routed submissions through individual real users rather than through the API — the clearest demonstration that rate limits govern channels, not actors.
- ★ Batching as an accidental buffer. Congressional offices absorbed AI-generated constituent mail relatively well because advocacy groups route through the "Communicating With Congress" API, which already batched high volume. Pre-existing infrastructure built for a different problem turned out to be the defence — worth noting as a counterexample to the assumption that every channel is equally exposed.
Structural limit: rate limits do nothing about qualitative flooding. Schmitz et al. find complexity increases in 90% of their 84 government cases versus volume increases in 60% — including a single German social-court letter running over 4,000 pages. A cap of six NIH applications does not stop each of the six from being longer and denser than any human would have written unaided. Most of this flood is not more submissions; it is heavier ones.
4. Disclosure & Transparency Requirements
Cheap to enact, and the category whose weakness the 2026 evidence exposes most clearly.
- ★★★ Disclosure fails when tested against detection. NeurIPS applied its lowest desk-reject threshold (Pangram ≥0.5) specifically to authors who declared no AI use or did not declare, catching 22 submissions. The declarations and the measurements disagreed, and the organizers concluded that "relying on author declarations is insufficient."
- ★★ The Common App treats AI-generated essay content as application fraud, authorizing account termination and revocation of admission — disclosure with a real penalty attached, which is rare. Institutions are separately piloting live writing samples, recorded video and audio statements. ⚠️ Policy wording from secondary sources.
- ★★ NSF requires disclosure of AI use in proposal preparation; non-disclosure is treated as misrepresentation. No caps, no originality declaration. ⚠️ The 26-1 PAPPG update was deferred pending OMB Uniform Guidance revisions, so this may tighten.
- ★★ Attestation fails the same way declarations do. PRiMER requires authors to attest that they did not use generative AI — an attestation that was violated in the very case that surfaced the problem. Two independent institutions (a small journal, a major conference) discovering the same thing is worth more than either alone.
- ★ 300+ AI-specific standing orders from US federal judges; some require identifying the specific tools used. In open source, Apache and Eclipse use commit-tag disclosure — the same instrument at a different scale.
- ★ California AI transparency laws (effective January 2026) require AI content providers to offer watermarks, latent disclosures, and detection tools — pushing the disclosure burden onto generators rather than submitters, which is the more promising design.
- ★ Nature, Elsevier, SAGE AI-disclosure requirements; COPE guidelines.
- ★ EU AI Act transparency mandates — ⚠️ high-risk obligations delayed 12–16 months (June 2026), so enforcement is further away than it was.
- ★ DDEX AI-credit metadata standard; California AB 412; Colorado State Fair rules.
Assessment: disclosure is not a response to flooding so much as a precondition for other responses, and it is only as strong as the verification behind it. Unenforced, it selects against honest declarers.
5. Outright Bans & Shutdowns
The nuclear option. Increasingly chosen not by cranks but by well-run projects that did the arithmetic.
★★★ curl ended its bug bounty (Daniel Stenberg, 26 Jan 2026). The best single exhibit in the whole project. The programme stopped 31 January 2026 having produced 87 confirmed vulnerabilities and over $100,000 paid; the confirmed-rate ran "north of 15%" in earlier years and "plummeted to below 5%" from 2025 — "Not even one in twenty was real." Reports are still accepted, via GitHub's private vulnerability reporting, with no bounty and no HackerOne. A functioning institution killed a functioning programme because triage cost exceeded signal value. That is the thesis in one paragraph. ⚠️ The commonly quoted "20 reports in the first 21 days of 2026, none real" is not in this post — it comes from secondary summaries; use the 15%→<5% collapse instead.
★★★ Wikipedia's G15 speedy-deletion criterion (adopted Aug 2025, updated by RfC 2026). Admins may immediately delete any page showing one of three unambiguous artifacts: user-directed communication ("Here is your Wikipedia article on…", "as a large language model", unfilled
[Birth Date]placeholders); non-existent or nonsensical references (invalid ISBN checksums, unrelated DOIs, citations predating the events they describe); or technical indicators (oaicite,attributionIndexblocks,turn0search0, markdown-fenced wikitext).This belongs in a different conceptual box from §1 detection, and that is the point. The policy explicitly bars using subjective stylistic "signs of AI writing" as the sole basis, because those may equally reflect human error or inexperience; borderline cases go to a noticeboard or a deletion discussion instead. Wikipedia bought a near-zero false-positive rate by only catching careless generation — the exact opposite trade from a classifier, and a genuine third option alongside "detect the style" and "verify the person." The 2026 RfC rationale is unusually candid about why: reasoning by analogy to the banned-user criterion, it accepts deleting some policy-compliant pages in order to "keep the cleanup workload manageable." An institution stating outright that review capacity, not content quality, is the binding constraint. ⚠️ Secondary coverage sometimes reports a broader March 2026 ban on LLM-generated article prose; that is unverified and does not appear in this policy.
★★ tldraw auto-closes all external pull requests; Ghostty zero-tolerance; Gentoo and NetBSD prohibit AI contributions. Against these, Kubernetes published guidance permitting AI-assisted contribution under accountability rules — the largest project in the world declining the ban.
★★ NIH: applications "substantially developed by AI" will not be considered — a ban nested inside the cap.
★ Clarkesworld temporarily closed all submissions (Feb 2023). The founding anecdote of this literature; keep it for narrative, not as current evidence.
★ Fulbright and Rhodes ban AI-generated content; Codeforces bans AI with detection; Hachette cancelled a contracted novel (Shy Girl, March 2026).
Note the asymmetry: every one of these closes a channel that was open because openness was cheap. The bug bounty, the unsolicited submission, the drive-by PR, the anonymous edit — these were gifts from a world where producing a plausible artifact was expensive. Bans are institutions revoking the gift.
6. Process Redesign (shifting to harder-to-fake interactions)
The most durable category, and the best-evidenced.
- ★★★ Universities are reviving oral examination at scale. Cornell, NYU Stern and Penn ran oral defences in spring 2026 midterms — 15-minute Socratic sessions, coding-while-talking, panel dialogues — after professors reported students submitting flawless homework and then failing to explain it. Phys.org, Aug 2026 is the honest version: oral exams have their own problems (anxiety, accessibility, bias in live judgment, and a cost that scales linearly with enrollment). Use this exhibit with its downside; the downside is the argument that redesign is expensive rather than free.
- ★★★ NeurIPS's audit-trail requirement is process redesign disguised as an appeals policy. Authors contesting a desk rejection must supply version history with a pre-AI checkpoint, a post-AI checkpoint, and analysis showing AI edits introduced no new substantive content. The organizers expect "this kind of audit trail will become a default." That shifts what is evaluated from the artifact to the provenance of the artifact — the deepest form of redesign in this taxonomy, because it is indifferent to how good generation gets.
- ★★ Blue books are back. Texas A&M, the University of Florida and UC Berkeley report surging demand for handwritten exam booklets over two years; from September 2026 UK universities move toward assessing process over output — drafts and revision memos.
- ★★ Service redesign as the durable government response. Schmitz et al.: structured interfaces, digital identity, and deterministic pipelines are the most durable answer to agentic flooding — and take years, which is why they predict demand suppression wins by default.
- ★★ Hiring shifts away from the résumé. Google, Cisco and McKinsey reinstated in-person interviews, and 72% of recruiting leaders now conduct them; 47% of hiring teams have updated interview techniques toward deeper assessment. ⚠️ Industry surveys.
- ★★ Redesigning the form rather than policing the submitter. Yale SOM researchers, having found that AI-edited complaints succeed more often (Yale Insights), propose two responses: integrate AI drafting into the complaint form itself so everyone gets it, or train staff to separate content from presentation when evaluating grievances. The only proposal in this taxonomy that answers a flood by levelling up access rather than restricting it — and the natural companion to the beneficial-flooding argument.
- ★ Process-based grading: drafts, revision memos, in-class writing.
- ★ ICML dual-track review; standardized recommendation-letter formats; structured planning-comment platforms; live VC demos over polished decks.
- ★ "Schools implementing redesign see 40% fewer integrity issues than detection-only approaches" — ⚠️ untraced to a primary. Do not cite until sourced; it is exactly the sort of too-convenient number this project should not launder.
7. Certification & Provenance
Marks of human origin. Still the weakest-adopted durable idea — but see the NeurIPS entry in §6, which is provenance arriving through the back door.
- ★★ Authors Guild "Human Authored" certification — $10/title, 3,000+ authors, 5,000+ titles.
- ★★ Documented method as "a form of truth defence." The Reuters Institute recommends journalists document every step of verification and cite each audiovisual source in ways others can replicate. Provenance applied to the investigator rather than the artifact — structurally the same move as NeurIPS's audit trail.
- ★ Reddit's visible "[APP]" label for registered bots.
- ★ DDEX AI-credit metadata; artist certificates and work-in-progress documentation for art contests.
- ★ Blockchain provenance for influencer marketing — ⚠️ vendor-promoted, no adoption evidence.
Why this stays low: certification asks the honest to bear a cost the dishonest can decline. It works only where the certificate gates something valuable, which is why the NeurIPS audit trail — where the "certificate" is the price of appealing a rejection — is more likely to stick than a voluntary badge.
8. Legislation & Regulation
Slowest, most authoritative, and currently going backwards in the US.
- ★★ FTC Consumer Review Rule (Oct 2024) bans AI-generated fake reviews; up to $51,744 per violation. The FTC sent warning letters to 10 companies over the 2025 holiday season signalling enforcement.
- ★★ Regulatory retreat is the live story. The FTC reopened and set aside the Rytr consent order (Dec 2025) under the Trump administration's AI Action Plan, calling the original order an undue burden on AI innovation. Meanwhile the EU AI Act's high-risk obligations were delayed 12–16 months (June 2026). Both major regimes moved away from enforcement during 2025–26.
- ★★ No federal rule constrains AI in political messaging ahead of the November 2026 midterms; the FEC remains deadlocked, leaving a patchwork of untested state laws. The FEC's Sept 2024 interpretive rule merely confirmed that existing fraudulent-misrepresentation provisions cover AI-generated communications — no disclaimers, no new rules.
- ★★ Comment Integrity and Management Act (2024, passed the House unanimously) — would require human verification of commenters, consolidated posting of mass comments, public disclosure of how many submissions were computer-generated, and uniform standards across agency dockets, with OMB guidance and a GAO report on identifying AI comments. The most fully specified legislative response anywhere in this taxonomy. Still not law. The executive branch moved separately: OMB directed OIRA to consider guidance on mass, computer-generated, and falsely attributed comments (Nov 2024).
- ★ AI Whistleblower Protection Act (S.1792); Ohio district AI-plan mandate; IRC §6673 penalties; FEC interpretive rule.
9. Financial Penalties & Sanctions
Reactive, but unusually well measured.
- ★★★ The Charlotin AI Hallucination Cases database (damiencharlotin.com/hallucinations) tracked roughly 1,490 decisions worldwide and more than 1,000 in the US as of May 2026, growing by several per day. Sanctions escalated from a $5,000 fine in 2023 to six-figure penalties, $15,000-per-attorney at the federal appellate level, and the first US bar suspensions tied to AI misuse by mid-2026. ⚠️ Exact current counts vary across secondary trackers (1,313 / 1,490 / 1,598 all circulate); cite the database itself with an access date rather than any single reported total. A live, public, growing dataset is a gift to this project — use it as the canonical measure of "courts are absorbing the cost."
- ★★ The cost lands on defendants, not just courts. Pro se cases that should cost ~2, 000todefendballoonto * *20,000–$70,000** under AI-enabled motion volume, and pro se litigants can file "four motions a week." Sanctions punish the filer; the defence cost is borne regardless of outcome — which is why penalties deter poorly here. ⚠️ Figures from practitioner reporting; trace before citing.
- ★★ NIH misconduct referrals, cost recovery, and grant termination for AI-content violations.
- ★ FTC $51,744/violation; Tax Court §6673 up to $25K; Section 512(f) suits over false DMCA claims (rarely used).
10. Industry Collaboration & Shared Infrastructure
Pooling detection and verification costs. Durable, underrated, quietly the model that scales.
- ★★★ STM Integrity Hub — 40 publishers, 125K papers/month screened. The clearest working instance of institutions solving a shared problem once instead of 40 times.
- ★★ NeurIPS's enterprise agreement with Pangram including zero data retention. A useful precedent: institutions can buy detection as infrastructure without handing over their members' work.
- ★★ Duplicate Submission Checker (12 publishers, 150+ journals); United2Act coalition.
- ★ Insurance Fraud Bureau (UK) and the 2024 Insurance Fraud Charter; chargeback consortium data sharing; NewsGuard + Pangram AI Content Farm detector.
11. Co-optation & Platform Integration
If you can't beat them, absorb them. Solves volume; changes what the interaction means.
- ★★★ AAAI 2026 formally embedded AI assistance across 22,977 reviews, and NeurIPS 2026 ran a randomized three-arm trial of AI-assisted reviewing (unassisted vs. two LLM-interaction conditions via OpenReview). ⚠️ AAAI figure from secondary coverage. The significance: peer review is running a controlled experiment on its own automation rather than only policing it — the first serious attempt in this taxonomy to find out whether co-optation actually works.
- ★★ NeurIPS bans reviewers from uploading papers to chatbots (confidentiality) while permitting AI for background research — co-optation and prohibition drawn along the confidentiality line rather than the quality line. A more defensible boundary than most.
- ★★ ~20% of US federal agencies now use ML or AI to process FOIA requests (CJR, Feb 2026); DoD and Dept. of Education have automated parts of intake. Agencies using AI to absorb requests that AI helped generate.
- ★★ Schmitz et al. find AI tools already deployed in processing in 25% of their 84 government flooding cases.
- ★★ Co-optation can fail expensively. The Commonwealth Bank of Australia reversed a decision to replace 45 customer-service roles with an AI "voice-bot" (August 2025) after the bot failed, call volumes surged and staff worked overtime. The only documented reversal of an automation-side response in this taxonomy, and the natural caution against assuming absorption is free. ⚠️ Trace to contemporaneous reporting before citing.
- ★ Upwork offers AI-drafted proposals as a platform feature; Change.org deepened AI integration; Hinge AI prompt feedback; UK "Extract" for planning objections. Freelancers publicly objected to Upwork's version — the losers of co-optation are visible in that case, which is unusual.
- ★ "68% of proposal teams use GenAI for RFPs"; "1 in 3 companies say AI will run hiring by 2026"; "95% of customer interactions AI-involved by 2025" — ⚠️ all vendor/consultancy surveys. Texture only.
12. Education & Behavioral Adaptation
Training humans to cope. Cheapest, weakest, but occasionally the only available move.
- ★★ Reuters Institute guidance that journalists treat unsolicited email as AI-generated until proven otherwise — a clean statement of the norm shift: the default presumption of good faith, which is what made tip lines work, formally withdrawn.
- ★ Behavior-based phishing training: 6× increase in reporting rates. ⚠️ Vendor-reported.
- ★ UC Berkeley's three-template syllabus policy (require AI / ban it / permit selective use).
- ★ NYT, Bellingcat, Guardian deepfake-detection workflows.
- ★ "Cold email extinction by 2027" — ⚠️ a blogger's prediction, repeatedly laundered. Drop it or label it as such.
13. Reduce the Expected Benefit
Make the flood not worth generating, by shrinking what a successful submission wins. Drawn from Schmitz, Hammond & Chan, whose four-class response framework treats it as a first-order strategy alongside adding friction.
- ★★★ Schmitz et al.'s risk matrix predicts exactly where this gets used. Near-term flooding risk is highest for financially attractive services whose demand was historically suppressed by friction and tacit knowledge — tax administration, court systems. Where the benefit is large and the friction was the only thing rationing it, institutions face a choice between capacity and entitlement.
- ★★ Narrower eligibility and smaller entitlements as the predicted default. Their forecast is Lindblom-style "muddling through": demand suppression wins because it is fast and proven, redesign loses because it takes years.
- ★ Australia considering reintroducing FOI fees after AI-generated request waves; US states opting for higher fees and longer extensions rather than reform (CJR).
Why this category matters more than its exhibit count suggests: it is the response with the worst distributive consequences and the least visibility. A cap is legible and contestable. A quietly narrowed eligibility rule is neither. If the project has a policy-relevant warning to give, it is probably here.
Summary: Which Responses Are Most Durable?
Detection cannot be ranked as a single thing; it splits in two, and the split does most of the work here.
Most durable:
- Process redesign — changing what is evaluated. Still first. The 2026 refinement: the most durable variant is not "make the task harder to fake" but "evaluate the provenance rather than the artifact" (NeurIPS audit trails, process-based grading, work-sample hiring). Cost is the binding constraint, not efficacy — oral exams scale linearly with enrollment.
- Shared infrastructure — the STM Hub model. Detection is expensive to do well; doing it once for 40 publishers is the only affordable path to doing it well at all.
- Identity verification — proving humanness beats detecting text. Constrained by adoption and by real privacy costs, not by efficacy.
Detection, split:
- Validated detection with published negative controls — moderately durable, and at NeurIPS decisive. It works when the institution treats the classifier as an instrument requiring calibration: negative controls on pre-2022 corpora, sensitivity analysis on windowing, conservative thresholds, corroborating evidence, and an appeals route. That is a demanding standard almost no institution currently meets.
- Unvalidated commodity detection — least durable, and being actively abandoned in education. Fails under domain and generator shift; cannot bear evidentiary weight against an individual.
Moderately durable:
- Volume caps — effective against quantitative flooding, useless against qualitative flooding, which is the more common form (90% vs. 60% in Schmitz et al.).
- Financial penalties — deterrent, reactive, and well measured thanks to the Charlotin database; but the cost falls on defendants regardless of who is sanctioned.
- Reduce expected benefit — durable in the sense that it works, at a cost borne by legitimate users.
Least durable:
- Disclosure requirements — falsified by measurement wherever both are applied (NeurIPS).
- Legislation — slow and, in the current US and EU environment, reversing.
Emerging and unproven:
- Prompt-injection traps (ICML) — effective, ethically contested, and dependent on a model limitation that will close.
- Provenance certification — voluntary versions have not scaled; the involuntary version (audit trail as the price of appeal) might.
- Co-optation — under genuine experiment (NeurIPS's three-arm trial, AAAI's 22,977 reviews) rather than mere assertion, and with at least one documented failure (Commonwealth Bank). Watch this.
Key Cross-Cutting Findings
No institution has fully solved it. Every response remains partial, and most remain reactive.
The detection arms race is not one race. The familiar claim — detection is always one generation behind — held for perplexity-based methods. Trained classifiers with published negative controls are a different instrument, and in 2026 they carried institutional weight for the first time. The honest formulation: detection is now good enough to measure a population and too contestable to convict an individual without corroboration. NeurIPS is the proof of both halves — it acted on the measurement, and it still built an appeals process.
Qualitative flooding is the bigger half, and most responses only touch the quantitative half. Complexity increases appear in 90% of Schmitz et al.'s cases, volume increases in 60%. Rate limits, caps, and CAPTCHAs address volume. A 4,000-page social-court letter defeats all of them.
The equity cost has moved, not vanished. From "whose prose reads as machine-like" (Liang et al. 2023, perplexity detectors) to "who can produce a documented audit trail" (NeurIPS 2026) and "who can afford the reintroduced fee or the in-person appearance" (Schmitz et al.). Each successive response is fairer than the last on its own terms and still sorts by resources. This is the finding most worth carrying into the paper.
Process redesign is underused because it is expensive, not because it is unknown. Institutions are not defaulting to detection out of ignorance; they are defaulting to it because oral exams cost faculty hours per student and service redesign takes years, while a classifier costs a licence fee.
Some floods are beneficial, and the beneficial ones are hardest to filter. AI-assisted CFPB complaints, FOIA requests, benefits appeals, and tax challenges democratize access. Schmitz et al. put this most sharply: agents genuinely reduce administrative burden and unlock legitimately suppressed demand, so the problem is capacity, not demand — and friction-based responses "lock this potential further out of reach." Note the tension with §13: reducing the expected benefit is precisely a decision to suppress demand rather than build capacity.
Physical presence and provenance are the two things AI cannot cheaply fake. In-person interviews, oral defences, blue books — and, increasingly, version history. Both are bets on evidence that exists outside the artifact.
The "both sides use AI" equilibrium is now measurable. FOIA (~20% of agencies process with AI), peer review (AAAI's 22,977 AI-assisted reviews; NeurIPS's randomized trial), RFPs, hiring, planning objections. In several domains this looks stable but is not a return to the status quo: both sides now spend more to reach the same decision.
The successful responses share a structure — measure first, act second, and publish the calibration. NeurIPS validated Pangram against a pre-ChatGPT control before rejecting anyone. curl counted its confirmation rate before killing the bounty. Wikipedia specified observable signatures before authorizing speedy deletion. The failures share the opposite structure: adopt a tool, act on its output, discover the error rate from the complaints.
Source quality
Read against primary sources: the NeurIPS desk rejections and methodology; the Pangram ICLR analysis; the curl bug-bounty closure; the NIH cap and AI policy; Curtin's detection disablement; FOIA volumes and agency AI use (CJR/Tow); Pudasaini et al.; Liang et al.; Spotify removals; the Verisk insurance-fraud survey; and Wikipedia's G15 policy text.
⚠️ Not verified — do not treat as load-bearing:
- "More than 50 universities have disabled Turnitin AI detection" — traces only to detector-marketing sites.
- "Schools implementing redesign see 40% fewer integrity issues" — no primary found.
- AAAI 2026's 22,977 AI-assisted reviews — secondary coverage only.
- The reported March 2026 Wikipedia ban on LLM-generated article prose — not in the speedy-deletion policy; locate it or drop it.
- LinkedIn's 99.65% proactive fake-account detection, and the Commonwealth Bank voice-bot reversal — both compelling, both currently second-hand.
- Pro se defence costs of 20, 000–70,000 and "four motions a week" — practitioner reporting.
- Clinical Orthopaedics 21-of-43 and the PRiMER attestation breach — from this project's own compilation, not yet traced.
- All vendor adoption percentages in §1, §11, §12 (65% of insurers, 62% of scholarship providers, 68% of schools, 68% of proposal teams, 95% of customer interactions).
- Charlotin database totals — the number moves; cite with an access date.
- Pangram's own false-positive rates — vendor self-reported, though third-party-validated per Pangram and independently exercised by NeurIPS's controls.
Known gap: no published Liang-style demographic false-positive audit of a frontier classifier. Until one exists, claims that modern detection is equitable are unsupported in the same way that claims it is inequitable are out of date.
A note on the supporting files. The five responses_*.md files contain useful per-type detail but carry no citations at all — roughly 1,500 lines with zero URLs between them. Anything promoted from them into this document is flagged above. They should be re-sourced before the project is published; until then, this file is the citable layer and they are working notes.