Model Autophagy Disorder (MAD)

A five-proposal research program investigating what happens to models trained on their own output, featuring pre-registered protocols and rigorous falsification arms.

What happens when an AI model is trained on its own output, and how early can we detect the resulting collapse? This is a five-proposal research program on Model Autophagy Disorder.

Three of the five proposals have working reference implementations, thoroughly tested via rigorous pre-registered protocols.

Preference-Induced Mode Collapse

I proved analytically that recursive autophagous DPO acts as the power map $p \mapsto p^\gamma/Z$ and verified it numerically to $1.4 \times 10^{-16}$. I measured recursive DPO losing effective diversity 11.3× faster than recursive SFT against a pre-registered floor of 2.5×.

Crucially, I built the control arm that could falsify my own hypothesis. An unbiased-judge control decomposed the collapse rate into 78.1% self-preference bias, 13.1% preference-estimator noise, and 8.8% neutral drift. Attributing the full effect to alignment without this control arm would have been wrong by an order of magnitude.

Early Detection and Verifiers

I developed mad_early for reference-free detection using tail statistics: frozen-reference coverage, Good-Turing missing mass, Hill tail index, Chao-Shen entropy, and more. To ensure validity, three invariants are enforced strictly in code:

  • The tail set is frozen at $t = 0$.
  • Entropy is always bias-corrected.
  • Sample size is held fixed across generations.

In vgsl (Verification-Guided Synthetic Loops), I measured homogenization under deterministic verifiers, treating the Wright-Fisher drift null as the baseline rather than zero. Generated code executes fully sandboxed with no network access and hard resource limits.

Pre-Registered Against Myself

Every protocol ships an amendment log recording every deviation, the reasoning behind it, and an explicit statement of whether the data had already been inspected. Post-hoc amendments automatically demote the affected test to exploratory, ensuring scientific rigor remains uncompromised.