Publications

See also: Google Scholar · OpenReview.

* denotes equal contribution.

Preprints

Sketch of the test loss versus the number of parameters p: the interpolation peak moves from p ~ n to p ~ nm, and after the peak the loss settles on a plateau above its earlier minimum.

Double Descent and Malign Overfitting in Diffusion Models

Raphaël Urfin*, Tony Bonnaire*, Giulio Biroli, Marc Mézard.

Preprint, 2026 Preprint

Conventional wisdom in deep learning says overparameterization is benign: models with more parameters $p$ than training samples $n$ still generalize well, and the test error follows a double-descent curve. Diffusion models seem to do the opposite — overparameterization leads to memorization. Studying the empirical score-matching loss with a fixed number $m$ of noised realizations per sample, we show that the interpolation peak shifts from the classical $p \sim n$ to $p \sim nm$, and that after the peak the test loss settles on a plateau higher than its earlier minimum — what we call malign overfitting. Since $m \gg 1$ in practice, real diffusion models sit on the rising branch before the peak, where the test loss already grows and the model starts memorizing. Overparameterization nonetheless remains beneficial with optimal regularization.

BibTeX
@misc{urfin2026doubledescentmalignoverfitting,
      title={Double Descent and Malign Overfitting in Diffusion Models},
      author={Raphaël Urfin and Tony Bonnaire and Giulio Biroli and Marc Mézard},
      year={2026},
      eprint={2609.26392},
      archivePrefix={arXiv},
      primaryClass={cs.LG},
      url={https://arxiv.org/abs/2609.26392},
}

2025

Schematic: two timescales τ_gen and τ_mem separate generalization from memorization during diffusion-model training, with τ_mem scaling linearly with dataset size n.

Why Diffusion Models Don’t Memorize: The Role of Implicit Dynamical Regularization in Training

Tony Bonnaire*, Raphaël Urfin*, Giulio Biroli, Marc Mézard.

NeurIPS 2025 Oral Best Paper

In this work, we investigate the role of the training dynamics in the transition from generalization to memorization in Diffusion Models. Through extensive experiments and theoretical analysis, we identify two distinct timescales: an early time $\tau_\mathrm{gen}$ at which models begin to generate high-quality samples, and a later time $\tau_\mathrm{mem}$ beyond which memorization emerges. Crucially, we find that $\tau_\mathrm{mem}$ increases linearly with the training set size $n$, while $\tau_\mathrm{gen}$ remains constant. This creates a growing window of training times with $n$ where models generalize effectively, despite showing strong memorization if training continues beyond it. It is only when $n$ becomes larger than a model-dependent threshold that overfitting disappears at infinite training times. These findings reveal a form of implicit dynamical regularization in the training dynamics, which allow to avoid memorization even in highly overparameterized settings.

BibTeX
@inproceedings{NEURIPS2025_ceb7f3cc,
 author = {Bonnaire, Tony and Urfin, Rapha\"{e}l and Biroli, Giulio and Mezard, Marc},
 booktitle = {Advances in Neural Information Processing Systems},
 editor = {D. Belgrave and C. Zhang and H. Lin and R. Pascanu and P. Koniusz and M. Ghassemi and N. Chen},
 pages = {141266--141286},
 publisher = {Curran Associates, Inc.},
 title = {Why Diffusion Models Don't Memorize: The Role of Implicit Dynamical Regularization in Training},
 url = {https://proceedings.neurips.cc/paper_files/paper/2025/file/ceb7f3cc876a6dcb15130a645b5a4507-Paper-Conference.pdf},
 volume = {38},
 year = {2025}
}