SiddharthAll posts
The Filter Is the Collapse: How AI Slop Compounds, and What It's Quietly Doing to Us

The Filter Is the Collapse: How AI Slop Compounds, and What It's Quietly Doing to Us

What you’ll learn
  • AI slop doesn't degrade by blurring output but by a plausibility filter that systematically deletes rare, novel, and dissenting content—the tails where breakthroughs live.
  • Experiments show recursion alone is stable, but keeping only "typical" outputs collapses diversity within generations, and any sustained synthetic share of new content compounds toward zero diversity.
  • The takeaway: the danger isn't machines talking to themselves, but our habit of treating their fluent output as authoritative, which quietly erodes the capacity for novelty.

We ran the loop ourselves — simulations, controls, dose curves — and found the danger isn't the machine talking to itself. It's the moment we decide its output is authoritative. This is the mathematics of that moment, and the six habits that keep you out of it.


The bug report civilization filed on itself

In May 2023, a media-integrity tracker counted 49 websites publishing machine-generated "news" with no humans in the loop. Nine months later there were more than 700. As of June 2026: 3,749, in sixteen languages, monetized by programmatic ads, dressed as local journalism www.newsguardtech.com www.voanews.com.

That's the small end. A web-scale study of 65,000 URLs — cross-checked with three independent detectors, false-positive rates under 2% — found that primarily-AI-written articles went from 2.2% of new English content in January 2020 to 35.9% within a year of the chatbot era, crossed 50% in early 2025, and have sat at rough parity ever since graphite.io. Peer-reviewed science is contaminated at the front door: excess-vocabulary analysis of 15 million biomedical abstracts puts a lower bound of 13.5% of 2024 abstracts as machine-processed — at least 200,000 papers a year, over 40% in some subfields www.science.org. In June 2026, for the first time in the web's history, bots generated more page requests than people www.business-standard.com. Two major dictionaries independently named slop their word of the year www.theguardian.com www.gadgetguy.com.au.

Most commentary stops at the inventory and calls it a garbage flood. That framing is wrong, and the truth is worse. The slop loop doesn't pass garbage through — it systematically deletes whatever is rare. And rarity is where novelty, dissent, and breakthrough live. The rest of this piece is the evidence and the arithmetic for that sentence, including an experiment we ran that isolates the exact step where the collapse happens. It isn't where you think.

The loop, precisely

Strip the branding and the slop economy is one machine with four gears. A person asks a model. The output reads authoritative — fluent, structured, confident — so it gets treated as ground truth: published as an article, pasted into a paper, cited as a source, or fed to another model as context. The corpus fills with machine text. The next model trains or retrieves on that corpus. Repeat.

Humanswrite · think · argueAI modelsgenerate outputAccepted asauthoritative?Publishedarticles · papersTHE CORPUSweb text + training dataDiscardedodd · rare · novelFresh human datathe regenerative inputpromptgeneratesyesfills the corpustrains the next modelnotails silently lostfresh signal
Fig 1 — The slop loop. The dashed paths are where the distribution loses its tails; the one accent path is the only input that regenerates them.

First, an honest result: the photocopy story is wrong

You've heard the folk version: AI feeding on AI output is "a photocopy of a photocopy" — each generation blurrier, collapse guaranteed. It's a satisfying story. Before claiming anything, we tested it, because a thesis built on a wrong mechanism will eventually embarrass you.

We built a toy world with eight "topics" — think of them as schools of thought, species, or market regimes — weighted from 30% down to 2%. The 2% topics are the rare, weird, valuable ones. Each generation, the model samples from its current beliefs and refits. Ten generations, five runs, fresh randomness each time.

Pure recursion was stable. All eight topics survived every generation. The fit to truth barely drifted. Independent theory work reaches the same conclusion: iterated retraining on accumulated data has stable fixed points arxiv.org. A machine reading its own words, left alone, does not melt down.

So if the loop is dangerous — and it is — the danger has to live somewhere else. It does. It lives in a step every one of us performs daily without noticing.

The bouncer: the one step that collapses everything

The real world doesn't run pure recursion. It runs curated recursion. We keep the plausible output and discard the weird one. We publish the fluent draft, delete the strange tangent, reward the answer that sounds right. Preference tuning formalizes exactly this: keep what scores well, train on that, repeat.

So we re-ran the experiment holding everything fixed — same toy world, same model, same data budget — and changed one rule: each generation, keep only the top 60% of outputs by how "typical" they look under the model's own beliefs. This is the smallest possible model of "I asked AI, the answer sounded authoritative, I kept it."

The result was not subtle.

Four recursive regimes compared across ten generations: control is stable, plausibility filter and truncated filter collapse to one mode, fresh data is stable
Fig 2 — The filter experiment. Four regimes, identical models, identical data budgets — only the selection rule differs. Recursion alone (gray) holds all 8 modes; keeping only the "plausible" 60% (indigo) collapses to a single spike; 20% fresh human data (teal) holds everything.

The first rare topic died in one generation. Eight topics became two within three generations. By generation ten, the model was a single narrow spike: its spread had shrunk to 10% of the truth, and the outer regions — the tails — held 0.2% of the mass they started with. Add a light trimming of the "least typical" material first (which is what low-temperature decoding and editorial cleanup do) and the collapse was total: one topic, no tails, in every single run. Meanwhile, the identical system fed 20% fresh human data per round held all eight topics and actually fit reality better than the untouched control.

Density snapshots across generations under the acceptance filter showing rare modes going extinct
Fig 3 — The same collapse, watched directly. Each row is the model's world at generations 0, 2, 4, 6, 8, 10 under the acceptance filter. Red arrows mark extinct rare modes. By generation 10, only the dominant peak remains.

Here's the mechanism, in plain language. A plausibility filter is a popularity contest judged by the current majority. Anything rare — a minority topic, an unusual finding, a genuinely novel claim — looks slightly "off" to a system tuned to what's common. So the filter strips it preferentially. The refit then shrinks its share further, which makes the next round's contest even more lopsided. It's a positive feedback loop on the selection rule itself, and it has a chilling formal property: conditioning on "accepted" is a variance-shrinking operator. Iterate it and the fixed point isn't the truth — it's the dominant mode. The weird stuff dies first, every time.

And it's a dial, not a cliff. In our dose-response sweep, keeping 97% of outputs — discarding a mere 3% per round as "not quite right" — still destroyed roughly 2.7 of the 8 topics over ten generations. Gentle curation, applied recursively, is just slower death. The published record agrees on the physics: recursive training on generated data causes defects from which the tails of the original distribution disappear first, across every model family tested www.nature.com arxiv.org, and self-consuming loops lose quality or diversity "in just a few generations" unless fresh real data enters each round arxiv.org.

Why the tails are everything

It's tempting to shrug: so the average gets a bit more average. But civilization doesn't run on the average. Scientific citations, technological breakthroughs, and successful strategies are heavy-tailed — the median paper, product, or plan contributes almost nothing; the progress all comes from the far right of the distribution. Breakthroughs are, by definition, things that looked implausible right up until they didn't.

Every mechanism in the slop loop — plausibility filtering, tail-trimming, homogenized ideation — is a tail-deletion operator. Run it long enough and you get a world that is perpetually 10% more polished and progressively less capable of producing anything that isn't already the mode. Call it median world: excellent answers to yesterday's questions, structural incapacity for the new ones. And the published work adds the sting: once the tails are gone from the training distribution, they stay gone — they can't be curated back, only re-earned from reality.

The arithmetic of compounding

Zoom from one model to the whole corpus and the loop becomes simple compound decay. Wherever machine text displaces fresh signal, diversity shrinks multiplicatively: each round, diversity gets multiplied by roughly (1 − 0.1 × slop share). Multiply enough times and there is no resting point where a slop-fed corpus stays diverse — the only fixed point of the equation is zero. The sole question is how fast you drain, and a half-life answers it:

Sustained synthetic share of new content5%10%20%40%
Generations for diversity to halve~139~69~35~17

Half-life = ln 2 ÷ (decay rate × slop share), from our compounding model (Fig 4). At today's ~50% share, the half-life is roughly a dozen loop generations.

Synthetic share growing logistically while diversity decays multiplicatively under three growth scenarios
Fig 4 — Corpus contamination compounding. The synthetic share grows logistically (left); diversity decays multiplicatively, with no non-zero resting point (right). Faster growth just moves the drain earlier.

Read that table as a statement about institutions, not algorithms. Any sustained synthetic share of the knowledge supply is lethal on a long enough horizon — the only variable is the half-life. And the plateau in the web data isn't safety; it's a slower leak at fifty percent.

Everyone ordering from the same menu

Now the demand side — what happens to our thinking when we ideate through the same model. Suppose each person's ideas come from their own hard-won distribution of experience, and the model's outputs come from one shared, somewhat narrower distribution. As adoption climbs, the pool of ideas in the room converges. The arithmetic (which our simulation matches exactly) is almost linear: idea diversity falls by roughly 0.52 × adoption rate.

Diversity ratio falling as AI adoption rises, simulation matching analytic curve
Fig 5 — The homogenization curve. At 25% adoption the idea space keeps 86% of its diversity; at 50%, 70%; at full adoption, 31%.

At a quarter of the room using the model, you've already lost 14% of the spread of thinking. At half, 30%. At full adoption, the room's ideas are 69% more alike than a world of individuals. The laboratory version is delightfully precise: 293 writers, randomly assigned AI assistance, produced stories that were individually more creative — and collectively 10.7% more similar to each other www.science.org. That's the trade in one sentence: everyone writes a better essay, and the world gets fewer different essays.

The cognitive ledger

The last gear of the loop is us. Three findings, each cleaner than the last:

  • The tutor study. Roughly a thousand high-school students, a full math curriculum segment. Students with an unguarded chatbot tutor improved 48% on practice problems — then scored 17% worse than the control group on the exam once the tool was removed. The tutor that carried them had eaten the skill it was supposed to teach. A guardrailed version of the same tutor produced +127% practice gains with no exam penalty www.pnas.org. The variable wasn't the AI. It was whether the human stayed in the loop.
  • The brain study. Over four months of essay-writing sessions, the group writing with LLM assistance showed the weakest neural connectivity of all groups — and 83% couldn't quote a single sentence from essays they had just written arxiv.org. The words were theirs to submit and not theirs to remember.
  • The workplace study. 319 knowledge workers, 936 documented examples of AI use in real tasks: the more confident people were in the AI, the less critical-thinking effort they deployed dl.acm.org. Confidence didn't add scrutiny; it subtracted it.

Put those into a model where critical-thinking capacity is a stock — eroding with disuse, growing with deliberate practice — and the trajectories separate fast. Behave as we currently do and the model says we lose roughly 39% of that capacity in thirty years. And it compounds reflexively: as judgment erodes, we get worse at detecting slop, so we outsource more, so judgment erodes further. Our simulation of that coupling shows the decline accelerating — modestly each year, relentlessly each decade.

Critical-thinking capital trajectories under guarded, default, and unchecked scenarios with and without reflexive coupling
Fig 6 — Critical-thinking capital, 30-year scenarios. The guarded path (gray) keeps deliberate practice; the default path (indigo) loses 39%; the dashed lines add the reflexive coupling — weaker judgment means less slop detection means more outsourcing.

This is cognitive debt, and it compounds exactly like technical debt — except the codebase is us.

The machine reading the machine

One more measurement, because it closes the loop physically. In 2025, AI crawlers averaged 4.2% of all page requests on one of the internet's largest infrastructure networks, while human traffic ended the year at 47% — already a minority of HTML requests once you count all automation blog.cloudflare.com. By June 2026, automated traffic crossed 57.5% and overtook humans outright www.business-standard.com.

The ratio that matters more: one frontier lab's crawler requested between 25,000 and 100,000 pages for every human visitor it sent back to the sites it read www.searchenginejournal.com. The machine extracts the web's grounding a hundred thousand times faster than it rewards it. Meanwhile, the human conversation that remains — the only truly fresh signal — is now fenced and licensed at premium rates, because scarcity does what scarcity does. We are eating the seed corn of cognition, and the price of seed corn is rising to confirm it.

It's governance, not fate

Here is the honest counterweight, and it's the reason to write this at all rather than shrug. Our own control run proves recursion is harmless when the loop is unbiased. Fresh data didn't just prevent collapse in our experiment — it produced the best fit of any regime. The dose curve is clean and monotone: 10% fresh human data per round restored the collapsed system to 3.7 of 8 topics; 20% to 6.0; 35% to 7.0; 50% to full recovery. The web's own contamination curve has plateaued rather than run away graphite.io.

Collapse is not a property of AI. It is a property of the loop's governance — and governance is a choice. Every number above that looks like a prophecy is actually a price tag on a behavior we can stop performing.

The protocol: six habits

  1. Verify, then integrate — never the reverse. An AI claim enters your work only with a primary source you actually opened. The experiment says the danger isn't the weird output; it's confidently accepting the plausible one.
  2. Never feed AI output to another AI as ground truth. That's the two-machine version of the acceptance filter — compounding with no human in the loop at all.
  3. Think first, prompt second. Draft your position before consulting the model, so it widens your distribution instead of replacing it. The brain data is unambiguous about the alternative.
  4. Interrogate, don't delegate. Use the model as a sparring partner — ask it to attack your reasoning, not to supply it. Adversarial use widens the search; oracle use narrows it.
  5. Institutions: buy fresh human data like the scarce input it now is. Expert time, field notes, original reporting, experiments — provenance-tracked. The 20%-fresh regime didn't just survive; it won. Human grounding isn't a cost center; it's the only thing in the system that regenerates.
  6. Keep the tails. High-temperature exploration, deliberately preserved dissent, archives of pre-machine text treated as conservation areas. Diversity is load-bearing.

The loop we built is efficient at exactly one thing: making tomorrow's plausible slightly more like today's. The mean is rising. The tails are dying. And only humans regenerate tails — which means the loop, as we've currently wired it, is quietly training us to stop doing the one thing it cannot do for itself.

That is not a reason to panic. It is a reason to hold the pen.

Image credits

Cover illustration
Generated for this article
AI-generated
0 comments
Siddharth
Siddharth

Thoughts and essays, published with Yokush. See more posts

Comments 0

Name & email required. Your email is never shown publicly.
No comments yet — be the first.