The African continent harbors the greatest human genetic diversity on Earth yet remains substantially underrepresented in global genomic reference datasets. This underrepresentation is particularly pronounced for Central-East Africa, and specifically for the Great Lakes region, which despite its exceptional demographic complexity and role as a convergence zone for Bantu, Nilotic, and pre- Bantu population movements, remains largely absent from major genomic panels. This study addresses this gap by characterizing the genomic profiles of two historically underrepresented African populations: Ethiopian individuals from the Amhara and Oromo groups (n=30, including five individuals from the Ethiopian-Kenyan border) and Rwandan individuals from Hutu, Tutsi, and Twa ethnic backgrounds (n=24, after quality control). Whole-genome sequencing data were generated for all groups and integrated with five African reference populations from the 1000 Genomes Project (Gambian, Mende, Esan, Yoruba, and Luhya; n=504) to construct a combined reference panel of 905,053 autosomal biallelic SNPs. Population structure was investigated using Principal Component Analysis (PCA) and model-based ancestry inference (ADMIXTURE), complemented by classification of mitochondrial and Y- chromosome haplogroups to reconstruct maternal and paternal histories. Given the lower and more heterogeneous sequencing coverage of the Rwandan samples, these individuals were projected onto the reference PCA eigenvectors rather than included in direct component computation. Results reveal two genetically distinct population histories. Ethiopian individuals display a complex ancestry shaped by ancient East African, Cushitic/Afroasiatic, and Eurasian components with mitochondrial haplogroups M1a and R0a and Y-chromosome haplogroup J1 supporting a documented migration from the Levant and Arabian Peninsula, and measurable internal structure distinguishing highland Amhara from Oromo/border-Kenya individuals. Rwandan individuals, by contrast, show a predominantly western Bantu autosomal and uniparental signature (mitochondrial haplogroup L and Y-chromosome E1b1a), clustering closely with West African Niger-Congo- speaking reference populations. Notably, no appreciable autosomal differentiation was observed among Hutu, Tutsi, and Twa individuals, supporting the view that these ethnic categories reflect socially and historically constructed classifications rather than deep biological divisions. Taken together, these findings contribute novel genomic data from two historically underrepresented African populations, refine current understanding of Ethiopian and Great Lakes population history, and underscore the importance of expanding African representation in genomic reference datasets.
Il continente africano custodisce la maggiore diversità genetica umana al mondo, ma rimane sostanzialmente sottorappresentato nei dataset genomici di riferimento a livello globale. Questa sottorappresentazione è particolarmente marcata per l'Africa centro-orientale, e in particolare per la regione dei Grandi Laghi, che nonostante la sua eccezionale complessità demografica e il suo ruolo di zona di convergenza per i movimenti di popolazioni Bantu, Nilotiche e pre-Bantu, resta in gran parte assente dai principali pannelli genomici. Questo studio affronta tale lacuna caratterizzando i profili genomici di due popolazioni africane storicamente sottorappresentate: individui etiopi appartenenti ai gruppi Amhara e Oromo (n=30, di cui cinque provenienti dal confine etiope-keniota) e individui ruandesi di origine Hutu, Tutsi e Twa (n=24, dopo il controllo di qualità). Per tutti i gruppi sono stati generati dati di sequenziamento dell'intero genoma, successivamente integrati con quelli di cinque popolazioni africane di riferimento del 1000 Genomes Project (Gambiani, Mende, Esan, Yoruba e Luhya; n=504), costruendo un pannello di riferimento combinato di 905.053 SNP autosomici bialleliici. La struttura di popolazione è stata indagata tramite Analisi delle Componenti Principali (PCA) e inferenza dell'ancestralità basata su modello (ADMIXTURE), integrate dalla classificazione degli aplogruppi mitocondriali e del cromosoma Y per ricostruire la storia materna e paterna. Data la scarsa ed eterogenea copertura di sequenziamento dei campioni ruandesi, questi individui sono stati proiettati sugli autovettori della PCA di riferimento anziché inclusi direttamente nel calcolo delle componenti. I risultati rivelano due storie di popolazione geneticamente distinte. Gli individui etiopi mostrano un'ancestralità complessa modellata da componenti est-africane antiche, cuscitiche/afroasiatiche ed eurasiatiche, con haplogruppi mitocondriali (M1a, R0a) e del cromosoma Y (J1) che avvalorano una documentata migrazione dal Levante e dalla Penisola Arabica, oltre a una struttura interna misurabile che distingue gli individui Amhara degli altipiani da quelli Oromo/del confine con il Kenya. Gli individui ruandesi, al contrario, mostrano una firma autosomica e uniparentale prevalentemente di origine Bantu occidentale (aplogruppo mitocondriale L, aplogruppo del cromosoma Y E1b1a), raggruppandosi strettamente con le popolazioni di riferimento dell'Africa occidentale di lingua Niger- Congo. È degno di nota che non sia stata osservata alcuna differenziazione autosomica apprezzabile tra individui Hutu, Tutsi e Twa, a sostegno dell'ipotesi che queste categorie etniche riflettano classificazioni socialmente e storicamente costruite piuttosto che profonde divisioni biologiche. Nel loro insieme, questi risultati contribuiscono con nuovi dati genomici relativi a due popolazioni africane storicamente sottorappresentate, affinano la comprensione attuale della storia di popolazione 2 etiope e dei Grandi Laghi, e sottolineano l'importanza di ampliare la rappresentazione africana nei dataset genomici di riferimento.
Sequenziamento dell'intero genoma e analisi della struttura della popolazione ruandese nel contesto genomico africano
LEGISTA, GIULIA
2025/2026
Abstract
The African continent harbors the greatest human genetic diversity on Earth yet remains substantially underrepresented in global genomic reference datasets. This underrepresentation is particularly pronounced for Central-East Africa, and specifically for the Great Lakes region, which despite its exceptional demographic complexity and role as a convergence zone for Bantu, Nilotic, and pre- Bantu population movements, remains largely absent from major genomic panels. This study addresses this gap by characterizing the genomic profiles of two historically underrepresented African populations: Ethiopian individuals from the Amhara and Oromo groups (n=30, including five individuals from the Ethiopian-Kenyan border) and Rwandan individuals from Hutu, Tutsi, and Twa ethnic backgrounds (n=24, after quality control). Whole-genome sequencing data were generated for all groups and integrated with five African reference populations from the 1000 Genomes Project (Gambian, Mende, Esan, Yoruba, and Luhya; n=504) to construct a combined reference panel of 905,053 autosomal biallelic SNPs. Population structure was investigated using Principal Component Analysis (PCA) and model-based ancestry inference (ADMIXTURE), complemented by classification of mitochondrial and Y- chromosome haplogroups to reconstruct maternal and paternal histories. Given the lower and more heterogeneous sequencing coverage of the Rwandan samples, these individuals were projected onto the reference PCA eigenvectors rather than included in direct component computation. Results reveal two genetically distinct population histories. Ethiopian individuals display a complex ancestry shaped by ancient East African, Cushitic/Afroasiatic, and Eurasian components with mitochondrial haplogroups M1a and R0a and Y-chromosome haplogroup J1 supporting a documented migration from the Levant and Arabian Peninsula, and measurable internal structure distinguishing highland Amhara from Oromo/border-Kenya individuals. Rwandan individuals, by contrast, show a predominantly western Bantu autosomal and uniparental signature (mitochondrial haplogroup L and Y-chromosome E1b1a), clustering closely with West African Niger-Congo- speaking reference populations. Notably, no appreciable autosomal differentiation was observed among Hutu, Tutsi, and Twa individuals, supporting the view that these ethnic categories reflect socially and historically constructed classifications rather than deep biological divisions. Taken together, these findings contribute novel genomic data from two historically underrepresented African populations, refine current understanding of Ethiopian and Great Lakes population history, and underscore the importance of expanding African representation in genomic reference datasets.| File | Dimensione | Formato | |
|---|---|---|---|
|
Tesi_Legista_Giulia.pdf
non disponibili
Dimensione
7.19 MB
Formato
Adobe PDF
|
7.19 MB | Adobe PDF | Richiedi una copia |
È consentito all'utente scaricare e condividere i documenti disponibili a testo pieno in UNITESI UNIPV nel rispetto della licenza Creative Commons del tipo CC BY NC ND.
Per maggiori informazioni e per verifiche sull'eventuale disponibilità del file scrivere a: [email protected].
https://hdl.handle.net/20.500.14239/35842