Sitemap
A list of all the posts and pages found on the site. For you robots out there, there is an XML version available for digesting as well.
Pages
Posts
Understanding diffusion models requires rethinking (again) generalization
Published:
We argue that the field should pivot from explaining why the diffusion models do not memorize to investigating what the model actually learns during pre-memorization phase.
portfolio
publications
Comptes Rendus Mathématique · 2022
More on lines in Euclidean Ramsey theory
Answers a question of Conlon and Fox, and independently of Arman and Tsaturian: there is an m and a red/blue colouring of Euclidean space in every dimension with no red copy of l_3 and no blue copy of l_m.
Conlon, D., & Wu, Y. H. (2023). More on lines in Euclidean Ramsey theory. Comptes Rendus. Mathématique, 361(G5), 897-901.
ICLR 2024
Implicit regularization of deep residual networks towards neural ODEs
Proves that a residual network initialized as a discretization of a neural ODE stays one throughout gradient-flow training, with convergence to a global minimum under a Polyak–Łojasiewicz condition.
Pierre Marion, Yu-Han Wu, Michael Eli Sander, & Gerard Biau (2024). Implicit regularization of deep residual networks towards neural ODEs. In The Twelfth International Conference on Learning Representations.
COLT 2025
Taking a Big Step: Large Learning Rates in Denoising Score Matching Prevent Memorization
Shows that the large learning rates used in practice implicitly regularize denoising score matching and keep training away from the memorizing empirical optimal score.
Wu, Y. H., Marion, P., Biau, G., & Boyer, C. (2025). Taking a Big Step: Large Learning Rates in Denoising Score Matching Prevent Memorization. Proceedings of Thirty Eighth Conference on Learning Theory, PMLR 291:5718-5756
ICML 2026
Optimal Stopping in Latent Diffusion Model
The last denoising steps of a latent diffusion model can degrade sample quality. This is intrinsic to the latent dimensionality reduction, and we give a principled account of when to stop.
Yu-Han Wu, Quentin Berthet, Gérard Biau, Claire Boyer, Romuald Elie & Pierre Marion (2025). Optimal Stopping in Latent Diffusion Model. arXiv preprint arXiv:2510.08409.
arXiv preprint · 2026
MIND: Monge Inception Distance for Generative Models Evaluation
A sliced-Wasserstein replacement for FID that is an order of magnitude more sample-efficient, two orders of magnitude faster to compute, and more robust to adversarial moment matching.
Quentin Berthet, Yu-Han Wu, Clément Crepy, Romuald Elie, Klaus Greff & Michael Eli Sander (2026). MIND: Monge Inception Distance for Generative Models Evaluation. arXiv:2605.06797.
arXiv preprint · 2026
Understanding diffusion models requires rethinking (again) generalization
A position paper arguing that generalization in diffusion models needs new theory: memorization and generalization are incompatible, so the question is what a model learns before it memorizes.
Pierre Marion and Yu-Han Wu (2026) Understanding diffusion models requires rethinking (again) generalization. arXiv:2605.06077.
arXiv preprint · 2026
DiffusionGemma
An experimental open-weight language model that generates text with discrete diffusion, obtained by fine-tuning Gemma 4, reaching about 1,500 output tokens per second on a single H100.
Taïga, DiffusionGemma Team Adrien Ali et al. “DiffusionGemma Technical Report.” (2026).
arXiv preprint · 2026
Kastor: An efficient fine-tuning strategy for generative emulation of PDE simulation
Turns a deterministic physics foundation model into a fast and accurate generative emulator, with a two-stage inference scheme and mean-prediction regularization, improving forecasts across The Well benchmark.
Guillaume Couairon, Alexis Jacq, Yu-Han Wu, Renu Singh, Yana Hasson, Quentin Berthet and Romuald Elie (2026) Kastor: An efficient fine-tuning strategy for generative emulation of PDE simulations. arXiv:2608.06107.
talks
Journées de Statistique 2025
Published:
Oral presentation of Taking a Big Step: Large Learning Rates in Denoising Score Matching Prevent Memorization.
COLT 2025
Published:
Oral presentation of Taking a Big Step: Large Learning Rates in Denoising Score Matching Prevent Memorization.
PriGM Workshop 2025
Published:
Oral presentation of Optimal Stopping in Latent Diffusion Models.
Lemanth 2026
Published:
Poster presentation of MIND: Monge Inception Distance for Generative Models Evaluation.
Optimisation et Apprentissage appliqués aux contenus numériques
Published:
Invited talk at USPN on Understanding diffusion models requires rethinking (again) generalization.
