Diffusion models have shown strong empirical performance in generative tasks, such as image generation, but it is still not fully understood how they generalize to unseen data, rather than simply memorizing the training set. This is an important question nowadays, with direct consequences for privacy and copyright. This thesis aims to study the generalization properties of diffusion models using tools from statistical physics and statistical learning theory, within a theoretically tractable Gaussian framework. We define generalization through a bound on the Kullback-Leibler (KL) divergence, which measures the discrepancy between the distribution learned by the empirical model and the true distribution from which the data are drawn. To control this divergence, we introduce a Leave-One-Out (LOO) influence function derived from the regularized sample covariance. We show that a generalized version of the KL divergence, which we call the excess KL divergence, is rigorously upper-bounded by the LOO influence density. Furthermore, by evaluating the thermodynamic limit via the Marchenko-Pastur spectral measure, we derive matching upper and lower bounds for the excess KL divergence, establishing that both quantities are of the same order across the whole range of sample complexities. We then connect these bounds to the convergence overlap factor q, first introduced in the Kadkhodaie and Guth framework and later formalized by Goldt and Maillard. This extends our analysis from a leave-one-out perturbation to a macroscopic comparison between models trained on disjoint data subsets, providing an exact analytical description of the memorization-to-generalization transition.
Nonostante gli ottimi risultati empirici ottenuti dai modelli di diffusione nei task gen- erativi (come, ad esempio, la generazione di immagini), non `e ancora del tutto chiaro come questi riescano a generalizzare su dati mai visti, piuttosto che limitarsi a mem- orizzare il training set. Si tratta di una questione cruciale, con dirette conseguenze, ad esempio, in materia di privacy e copyright. Questo lavoro affronta le propriet`a di generalizzazione dei modelli di diffusione attraverso tecniche di fisica statistica e di sta- tistical learning theory, all’interno di un framework Gaussiano trattabile analiticamente. Definiamo la generalizzazione attraverso un limite sulla divergenza di Kullback-Leibler (KL), che quantifica la discrepanza tra la distribuzione appresa dal modello empirico e la vera distribuzione generatrice dei dati. Per controllare tale divergenza, introduciamo una funzione di influenza Leave-One-Out (LOO) derivata dalla matrice di covarianza campionaria regolarizzata. Mostriamo che una versione generalizzata della divergenza KL, denominata excess Kullback-Leibler divergence, `e rigorosamente limitata superior- mente dalla densit`a di influenza LOO. Inoltre, valutando il limite termodinamico per n e d, tramite la misura spettrale di Marchenko-Pastur, ricaviamo dei limiti (upper e lower bounds) per la KL, dimostrando che divergenza e influenza sono dello stesso ordine di grandezza per qualsiasi complessit`a campionaria. Infine, colleghiamo questi risultati al fattore di overlap q, introdotto nel framework di Kadkhodaie e Guth e formalizzato successivamente da Goldt e Maillard. Questo passaggio estende l’analisi da un con- testo di perturbazione leave-one-out al confronto macroscopico tra modelli addestrati su sottoinsiemi di dati disgiunti, fornendo una descrizione analitica esatta della transizione da memorizzazione a generalizzazione.
Funzioni di Influenza e Convergenza nei Modelli Generativi
BOMBARI, CECILIA
2025/2026
Abstract
Diffusion models have shown strong empirical performance in generative tasks, such as image generation, but it is still not fully understood how they generalize to unseen data, rather than simply memorizing the training set. This is an important question nowadays, with direct consequences for privacy and copyright. This thesis aims to study the generalization properties of diffusion models using tools from statistical physics and statistical learning theory, within a theoretically tractable Gaussian framework. We define generalization through a bound on the Kullback-Leibler (KL) divergence, which measures the discrepancy between the distribution learned by the empirical model and the true distribution from which the data are drawn. To control this divergence, we introduce a Leave-One-Out (LOO) influence function derived from the regularized sample covariance. We show that a generalized version of the KL divergence, which we call the excess KL divergence, is rigorously upper-bounded by the LOO influence density. Furthermore, by evaluating the thermodynamic limit via the Marchenko-Pastur spectral measure, we derive matching upper and lower bounds for the excess KL divergence, establishing that both quantities are of the same order across the whole range of sample complexities. We then connect these bounds to the convergence overlap factor q, first introduced in the Kadkhodaie and Guth framework and later formalized by Goldt and Maillard. This extends our analysis from a leave-one-out perturbation to a macroscopic comparison between models trained on disjoint data subsets, providing an exact analytical description of the memorization-to-generalization transition.| File | Dimensione | Formato | |
|---|---|---|---|
|
master_thesis_bombari-1.pdf
accesso aperto
Descrizione: Documento di tesi magistrale
Dimensione
5.17 MB
Formato
Adobe PDF
|
5.17 MB | Adobe PDF | Visualizza/Apri |
È consentito all'utente scaricare e condividere i documenti disponibili a testo pieno in UNITESI UNIPV nel rispetto della licenza Creative Commons del tipo CC BY NC ND.
Per maggiori informazioni e per verifiche sull'eventuale disponibilità del file scrivere a: [email protected].
https://hdl.handle.net/20.500.14239/36521