In this thesis, we will see how the theory of gradient flows and optimal transport can be used to describe neural network training. The classical formulation of training can be seen as a gradient flow with respect to the Euclidean norm, but the non-convexity of the optimization landscape complicates the analysis of convergence to global minimizers. Here, we consider the case of shallow neural networks, reformulating the training in terms of probability measures. Interpreting the set of parameters of N neurons as an interactive particle system, we prove that there exists a unique particle gradient flow for a regularized loss functional. We also prove that, when N goes to infinity, the particle gradient flow converges to the unique Wasserstein gradient flow for the loss functional. This result is called many-particle limit and allows us to interpret the training of an infinite-width shallow neural network as a flow governed by a continuity equation. We recall the analytical conditions established by Chizat and Bach (2018) under which the Wasserstein gradient flow is guaranteed to converge to a global minimizer. Considering a finite number of neurons, we show that the curve of empirical measures associated with a particle gradient flow is already a Wasserstein gradient flow. To explore visible differences between these two dynamics, which are driven by different metrics, we conduct a series of numerical experiments using a neural network with a ReLU activation, discretizing the Wasserstein gradient flow via the JKO scheme and the Sinkhorn algorithm, and the particle gradient flow via the batch gradient descent algorithm. The experiments show that the difference between the dynamics is minimal according to the many-particle limit.
In questa tesi, vedremo come la teoria dei flussi gradiente e del trasporto ottimo possa essere utilizzata per descrivere l'addestramento delle reti neurali. La formulazione classica dell'addestramento può essere vista come un gradient flow rispetto alla norma euclidea, ma la non convessità del panorama di ottimizzazione complica l'analisi della convergenza verso i minimizzatori globali. Qui, consideriamo il caso delle reti neurali shallow, riformulando l'addestramento in termini di misure di probabilità. Interpretando l'insieme dei parametri di N neuroni come un sistema di particelle interagenti, dimostriamo che esiste un unico particle gradient flow per un funzionale di costo regolarizzato. Dimostriamo inoltre che, quando N tende all'infinito, il particle gradient flow converge all'unico Wasserstein gradient flow per il funzionale di costo. Questo risultato è chiamato many-particle limit e ci permette di interpretare l'addestramento di una rete neurale shallow di larghezza infinita come un flusso governato da un'equazione di continuità. Richiamiamo le condizioni analitiche stabilite da Chizat e Bach (2018) sotto le quali è garantito che il Wasserstein gradient flow converga a un minimizzatore globale. Considerando un numero finito di neuroni, mostriamo che la curva di misure empiriche associata a un particle gradient flow è già un Wasserstein gradient flow. Per esplorare le differenze visibili tra queste due dinamiche, che sono guidate da metriche differenti, conduciamo una serie di esperimenti numerici utilizzando una rete neurale con attivazione ReLU, discretizzando il Wasserstein gradient flow tramite lo schema JKO e l'algoritmo di Sinkhorn, e il particle gradient flow tramite l'algoritmo di batch gradient descent. Gli esperimenti mostrano che la differenza tra le dinamiche è minima in accordo con il many-particle limit.
Flussi Gradiente e Trasporto Ottimo nell'Addestramento delle Reti Neurali
BUCCIERI, EDWARD FRANCISCO
2025/2026
Abstract
In this thesis, we will see how the theory of gradient flows and optimal transport can be used to describe neural network training. The classical formulation of training can be seen as a gradient flow with respect to the Euclidean norm, but the non-convexity of the optimization landscape complicates the analysis of convergence to global minimizers. Here, we consider the case of shallow neural networks, reformulating the training in terms of probability measures. Interpreting the set of parameters of N neurons as an interactive particle system, we prove that there exists a unique particle gradient flow for a regularized loss functional. We also prove that, when N goes to infinity, the particle gradient flow converges to the unique Wasserstein gradient flow for the loss functional. This result is called many-particle limit and allows us to interpret the training of an infinite-width shallow neural network as a flow governed by a continuity equation. We recall the analytical conditions established by Chizat and Bach (2018) under which the Wasserstein gradient flow is guaranteed to converge to a global minimizer. Considering a finite number of neurons, we show that the curve of empirical measures associated with a particle gradient flow is already a Wasserstein gradient flow. To explore visible differences between these two dynamics, which are driven by different metrics, we conduct a series of numerical experiments using a neural network with a ReLU activation, discretizing the Wasserstein gradient flow via the JKO scheme and the Sinkhorn algorithm, and the particle gradient flow via the batch gradient descent algorithm. The experiments show that the difference between the dynamics is minimal according to the many-particle limit.| File | Dimensione | Formato | |
|---|---|---|---|
|
Tesi_Magistrale_Buccieri_547273 (1).pdf
accesso aperto
Dimensione
13.19 MB
Formato
Adobe PDF
|
13.19 MB | Adobe PDF | Visualizza/Apri |
È consentito all'utente scaricare e condividere i documenti disponibili a testo pieno in UNITESI UNIPV nel rispetto della licenza Creative Commons del tipo CC BY NC ND.
Per maggiori informazioni e per verifiche sull'eventuale disponibilità del file scrivere a: [email protected].
https://hdl.handle.net/20.500.14239/36103