This thesis examines whether machine-learning-based expected return estimates can improve out-of-sample portfolio performance compared with the traditional Historical Mean benchmark. It also studies whether the model with the best financial performance is also the most compliant from a SAFE AI perspective. The empirical analysis uses a stock universe derived from the S&P 500, based on Bloomberg data, from January 2010 to December 2025. The out-of-sample period covers 2023 to 2025. Four return estimation methods are compared: Historical Mean, Ridge Regression, Neural Network and XGBoost. All models are evaluated within the same constrained portfolio optimization process, using the same covariance estimator, constraints, rebalancing rule and transaction cost assumptions. The SAFE AI assessment considers robustness, accuracy, fairness and explainability. Since the predictions concern securities rather than individuals, fairness is measured across GICS sectors. The SAFE AI dimensions are then combined into an integrated compliance score and linked to portfolio performance through the SAFE-Performance Frontier. The results indicate that XGBoost delivers the strongest out-of-sample risk-adjusted performance and the highest SAFE AI compliance score, although its advantage over the Historical Mean benchmark is reduced after transaction costs due to higher turnover.

Questa tesi analizza se le stime dei rendimenti attesi ottenute tramite modelli di machine learning possano migliorare la performance out-of-sample di portafoglio rispetto al benchmark tradizionale della media storica. Inoltre, lo studio verifica se il modello con la migliore performance finanziaria sia anche quello più conforme dal punto di vista SAFE AI. L’analisi empirica utilizza un universo di titoli derivato dall’S&P 500, costruito con dati Bloomberg, da gennaio 2010 a dicembre 2025. Il periodo out-of-sample considerato va dal 2023 al 2025. Vengono confrontati quattro metodi di stima dei rendimenti: media storica, Ridge Regression, Neural Network e XGBoost. Tutti i modelli sono valutati nello stesso processo di ottimizzazione vincolata di portafoglio, usando lo stesso stimatore della covarianza, gli stessi vincoli, la stessa regola di ribilanciamento e la stessa struttura dei costi di transazione. La valutazione SAFE AI considera robustezza, accuratezza, fairness e spiegabilità. Poiché le previsioni riguardano titoli finanziari e non individui, la fairness viene misurata tra i settori GICS. Le dimensioni SAFE AI sono poi sintetizzate in un compliance score integrato e collegate alla performance attraverso la SAFE-Performance Frontier. I risultati indicano che XGBoost ottiene la migliore performance out-of-sample corretta per il rischio e il compliance score SAFE AI più elevato, sebbene il suo vantaggio rispetto alla media storica si riduca dopo i costi di transazione a causa del maggiore turnover.

Machine learning e ottimizzazione di portafoglio nel framework SAFE AI: uno studio empirico sulla stima dei rendimenti attesi

VATA, ANILA
2025/2026

Abstract

This thesis examines whether machine-learning-based expected return estimates can improve out-of-sample portfolio performance compared with the traditional Historical Mean benchmark. It also studies whether the model with the best financial performance is also the most compliant from a SAFE AI perspective. The empirical analysis uses a stock universe derived from the S&P 500, based on Bloomberg data, from January 2010 to December 2025. The out-of-sample period covers 2023 to 2025. Four return estimation methods are compared: Historical Mean, Ridge Regression, Neural Network and XGBoost. All models are evaluated within the same constrained portfolio optimization process, using the same covariance estimator, constraints, rebalancing rule and transaction cost assumptions. The SAFE AI assessment considers robustness, accuracy, fairness and explainability. Since the predictions concern securities rather than individuals, fairness is measured across GICS sectors. The SAFE AI dimensions are then combined into an integrated compliance score and linked to portfolio performance through the SAFE-Performance Frontier. The results indicate that XGBoost delivers the strongest out-of-sample risk-adjusted performance and the highest SAFE AI compliance score, although its advantage over the Historical Mean benchmark is reduced after transaction costs due to higher turnover.
2025
Machine Learning in Portfolio Optimization under the SAFE AI Framework: An Empirical Study of Expected Return Estimation
Questa tesi analizza se le stime dei rendimenti attesi ottenute tramite modelli di machine learning possano migliorare la performance out-of-sample di portafoglio rispetto al benchmark tradizionale della media storica. Inoltre, lo studio verifica se il modello con la migliore performance finanziaria sia anche quello più conforme dal punto di vista SAFE AI. L’analisi empirica utilizza un universo di titoli derivato dall’S&P 500, costruito con dati Bloomberg, da gennaio 2010 a dicembre 2025. Il periodo out-of-sample considerato va dal 2023 al 2025. Vengono confrontati quattro metodi di stima dei rendimenti: media storica, Ridge Regression, Neural Network e XGBoost. Tutti i modelli sono valutati nello stesso processo di ottimizzazione vincolata di portafoglio, usando lo stesso stimatore della covarianza, gli stessi vincoli, la stessa regola di ribilanciamento e la stessa struttura dei costi di transazione. La valutazione SAFE AI considera robustezza, accuratezza, fairness e spiegabilità. Poiché le previsioni riguardano titoli finanziari e non individui, la fairness viene misurata tra i settori GICS. Le dimensioni SAFE AI sono poi sintetizzate in un compliance score integrato e collegate alla performance attraverso la SAFE-Performance Frontier. I risultati indicano che XGBoost ottiene la migliore performance out-of-sample corretta per il rischio e il compliance score SAFE AI più elevato, sebbene il suo vantaggio rispetto alla media storica si riduca dopo i costi di transazione a causa del maggiore turnover.
File in questo prodotto:
File Dimensione Formato  
Thesis_Anila_Vata.pdf

embargo fino al 28/07/2027

Dimensione 3.64 MB
Formato Adobe PDF
3.64 MB Adobe PDF   Richiedi una copia

È consentito all'utente scaricare e condividere i documenti disponibili a testo pieno in UNITESI UNIPV nel rispetto della licenza Creative Commons del tipo CC BY NC ND.
Per maggiori informazioni e per verifiche sull'eventuale disponibilità del file scrivere a: [email protected].

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/20.500.14239/35989