Double Descent and Malign Overfitting in Diffusion Models
Preprint, 2026 Preprint
Conventional wisdom in deep learning says overparameterization is benign: models with more parameters $p$ than training samples $n$ still generalize well, and the test error follows a double-descent curve. Diffusion models seem to do the opposite — overparameterization leads to memorization. Studying the empirical score-matching loss with a fixed number $m$ of noised realizations per sample, we show that the interpolation peak shifts from the classical $p \sim n$ to $p \sim nm$, and that after the peak the test loss settles on a plateau higher than its earlier minimum — what we call malign overfitting. Since $m \gg 1$ in practice, real diffusion models sit on the rising branch before the peak, where the test loss already grows and the model starts memorizing. Overparameterization nonetheless remains beneficial with optimal regularization.
BibTeX
@misc{urfin2026doubledescentmalignoverfitting,
title={Double Descent and Malign Overfitting in Diffusion Models},
author={Raphaël Urfin and Tony Bonnaire and Giulio Biroli and Marc Mézard},
year={2026},
eprint={2609.26392},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2609.26392},
}