Publications

You can also find my articles on my Google Scholar profile.

2026

arXiv preprint · 2026

Kastor: An efficient fine-tuning strategy for generative emulation of PDE simulation

Guillaume Couairon, Alexis Jacq, Yu-Han Wu, Renu Singh, Yana Hasson, Quentin Berthet, Romuald Elie

Turns a deterministic physics foundation model into a fast and accurate generative emulator, with a two-stage inference scheme and mean-prediction regularization, improving forecasts across The Well benchmark.

Guillaume Couairon, Alexis Jacq, Yu-Han Wu, Renu Singh, Yana Hasson, Quentin Berthet and Romuald Elie (2026) Kastor: An efficient fine-tuning strategy for generative emulation of PDE simulations. arXiv:2608.06107.

Paper arXiv Details

arXiv preprint · 2026

DiffusionGemma

DiffusionGemma Team, Google DeepMind

An experimental open-weight language model that generates text with discrete diffusion, obtained by fine-tuning Gemma 4, reaching about 1,500 output tokens per second on a single H100.

Taïga, DiffusionGemma Team Adrien Ali et al. “DiffusionGemma Technical Report.” (2026).

Paper arXiv Details

arXiv preprint · 2026

Understanding diffusion models requires rethinking (again) generalization

Pierre Marion*, Yu-Han Wu* (* equal contribution)

A position paper arguing that generalization in diffusion models needs new theory: memorization and generalization are incompatible, so the question is what a model learns before it memorizes.

Pierre Marion and Yu-Han Wu (2026) Understanding diffusion models requires rethinking (again) generalization. arXiv:2605.06077.

Paper arXiv Details

arXiv preprint · 2026

MIND: Monge Inception Distance for Generative Models Evaluation

Quentin Berthet, Yu-Han Wu, Clément Crepy, Romuald Elie, Klaus Greff, Michael E. Sander

A sliced-Wasserstein replacement for FID that is an order of magnitude more sample-efficient, two orders of magnitude faster to compute, and more robust to adversarial moment matching.

Quentin Berthet, Yu-Han Wu, Clément Crepy, Romuald Elie, Klaus Greff & Michael Eli Sander (2026). MIND: Monge Inception Distance for Generative Models Evaluation. arXiv:2605.06797.

Paper arXiv Poster Details

2025

ICML 2026

Optimal Stopping in Latent Diffusion Model

Yu-Han Wu, Quentin Berthet, Gérard Biau, Claire Boyer, Romuald Elie, Pierre Marion

The last denoising steps of a latent diffusion model can degrade sample quality. This is intrinsic to the latent dimensionality reduction, and we give a principled account of when to stop.

Yu-Han Wu, Quentin Berthet, Gérard Biau, Claire Boyer, Romuald Elie & Pierre Marion (2025). Optimal Stopping in Latent Diffusion Model. arXiv preprint arXiv:2510.08409.

Paper arXiv Slides Poster Details

COLT 2025

Taking a Big Step: Large Learning Rates in Denoising Score Matching Prevent Memorization

Yu-Han Wu, Pierre Marion, Gérard Biau, Claire Boyer

Shows that the large learning rates used in practice implicitly regularize denoising score matching and keep training away from the memorizing empirical optimal score.

Wu, Y. H., Marion, P., Biau, G., & Boyer, C. (2025). Taking a Big Step: Large Learning Rates in Denoising Score Matching Prevent Memorization. Proceedings of Thirty Eighth Conference on Learning Theory, PMLR 291:5718-5756

Paper arXiv Slides Details

2023

ICLR 2024

Implicit regularization of deep residual networks towards neural ODEs

Pierre Marion*, Yu-Han Wu*, Michael E. Sander, Gérard Biau (* equal contribution)

Proves that a residual network initialized as a discretization of a neural ODE stays one throughout gradient-flow training, with convergence to a global minimum under a Polyak–Łojasiewicz condition.

Pierre Marion, Yu-Han Wu, Michael Eli Sander, & Gerard Biau (2024). Implicit regularization of deep residual networks towards neural ODEs. In The Twelfth International Conference on Learning Representations.

Paper arXiv Details

2022

Comptes Rendus Mathématique · 2022

More on lines in Euclidean Ramsey theory

David Conlon, Yu-Han Wu

Answers a question of Conlon and Fox, and independently of Arman and Tsaturian: there is an m and a red/blue colouring of Euclidean space in every dimension with no red copy of l_3 and no blue copy of l_m.

Conlon, D., & Wu, Y. H. (2023). More on lines in Euclidean Ramsey theory. Comptes Rendus. Mathématique, 361(G5), 897-901.

Paper arXiv Details