Alternative splicing of transcription factor genes generates protein isoforms that differ substantially from their reference counterparts in DNA-binding specificity and transcriptional output, yet the structural basis for this functional divergence at the protein-DNA interface remains poorly characterized. We present a systematic structural comparison of 693 isoforms derived from 246 transcription factor genes, using three complementary deep learning frameworks, AlphaFold 3, RoseTTAFold2NA, and Boltz-2, to model and compare the predicted three-dimensional structures of reference and alternative isoform-DNA complexes. We show that structural confidence metrics from all three platforms are sensitive to isoform-specific sequence differences located outside the DNA-binding domain, and that alternative isoforms retaining a fully intact DNA-binding domain consistently show lower predicted inter-domain positional error than their reference counterparts across all three platforms. This pattern of structural stabilization is consistent with an autoinhibitory model in which flanking intrinsically disordered regions in reference isoforms physically constrain the DNA-binding domain and limit its conformational freedom, and their removal through alternative splicing releases this constraint. Analysis of eight isoforms selected through experimental filtering of discordant protein-DNA interaction data shows that structural stabilization signals are most consistent and interpretable when all three platforms agree, with DLX4-2 showing the strongest cross-platform evidence supported by experimental data, and PRRX1-3 representing the most coherent novel computational candidate. Receiver operating characteristic analysis across 79 unique isoforms demonstrates that no individual structural metric reliably predicts experimental binding outcomes when the models are applied without training on binding data, though a combined logistic regression model integrating metrics from all three platforms achieves an area under the curve of 0.730. These results provide a benchmarked multi-platform computational framework for characterizing structural differences between transcription factor isoforms and establish a performance baseline for future approaches that incorporate experimental binding data into model training.
Lo splicing alternativo dei geni dei fattori di trascrizione genera isoforme proteiche che possono differire in modo sostanziale dalle controparti di riferimento per specificità di legame al DNA e attività trascrizionale, tuttavia la base strutturale di questa divergenza funzionale all'interfaccia proteina-DNA rimane ancora poco caratterizzata. Presentiamo un confronto strutturale sistematico di 693 isoforme derivate da 246 geni di fattori di trascrizione, utilizzando tre framework di deep learning complementari, AlphaFold 3, RoseTTAFold2NA e Boltz-2, per modellare e confrontare le strutture tridimensionali predette dei complessi isoforma-DNA di riferimento e alternativi. Dimostriamo che le metriche di confidenza strutturale di tutte e tre le piattaforme sono sensibili alle differenze di sequenza specifiche delle isoforme localizzate al di fuori del dominio di legame al DNA, e che le isoforme alternative che conservano un dominio di legame al DNA completamente intatto mostrano in modo consistente un errore posizionale inter-dominio predetto inferiore rispetto alle controparti di riferimento in tutte e tre le piattaforme. Questo schema di stabilizzazione strutturale è coerente con un modello autoinibitori in cui le regioni intrinsecamente disordinate fiancheggianti nelle isoforme di riferimento vincolano fisicamente il dominio di legame al DNA e ne limitano la libertà conformazionale, e la loro rimozione attraverso lo splicing alternativo libera questo vincolo. L'analisi di otto isoforme selezionate attraverso il filtraggio sperimentale di dati discordanti di interazione proteina-DNA mostra che i segnali di stabilizzazione strutturale sono più consistenti e interpretabili quando tutte e tre le piattaforme concordano, con DLX4-2 che mostra le evidenze più forti supportate da dati sperimentali, e PRRX1-3 che rappresenta il candidato computazionale novel più coerente. L'analisi della curva ROC su 79 isoforme uniche dimostra che nessuna metrica strutturale individuale predice in modo affidabile gli esiti di legame sperimentali quando i modelli vengono applicati senza addestramento su dati di legame, sebbene un modello di regressione logistica combinato che integra le metriche di tutte e tre le piattaforme raggiunga un'area sotto la curva di 0.730. Questi risultati forniscono un framework computazionale multi-piattaforma validato tramite benchmarking per caratterizzare le differenze strutturali tra le isoforme dei fattori di trascrizione e stabiliscono una base di riferimento per futuri approcci che incorporano dati di legame sperimentali nell'addestramento dei modelli.
Previsione computazionale delle interazioni tra isoforme di fattori di trascrizione (TF) e DNA mediante RoseTTAFold2NA, AlphaFold 3 e Boltz-2
SOLEIMANI JEVINANI, SARA
2025/2026
Abstract
Alternative splicing of transcription factor genes generates protein isoforms that differ substantially from their reference counterparts in DNA-binding specificity and transcriptional output, yet the structural basis for this functional divergence at the protein-DNA interface remains poorly characterized. We present a systematic structural comparison of 693 isoforms derived from 246 transcription factor genes, using three complementary deep learning frameworks, AlphaFold 3, RoseTTAFold2NA, and Boltz-2, to model and compare the predicted three-dimensional structures of reference and alternative isoform-DNA complexes. We show that structural confidence metrics from all three platforms are sensitive to isoform-specific sequence differences located outside the DNA-binding domain, and that alternative isoforms retaining a fully intact DNA-binding domain consistently show lower predicted inter-domain positional error than their reference counterparts across all three platforms. This pattern of structural stabilization is consistent with an autoinhibitory model in which flanking intrinsically disordered regions in reference isoforms physically constrain the DNA-binding domain and limit its conformational freedom, and their removal through alternative splicing releases this constraint. Analysis of eight isoforms selected through experimental filtering of discordant protein-DNA interaction data shows that structural stabilization signals are most consistent and interpretable when all three platforms agree, with DLX4-2 showing the strongest cross-platform evidence supported by experimental data, and PRRX1-3 representing the most coherent novel computational candidate. Receiver operating characteristic analysis across 79 unique isoforms demonstrates that no individual structural metric reliably predicts experimental binding outcomes when the models are applied without training on binding data, though a combined logistic regression model integrating metrics from all three platforms achieves an area under the curve of 0.730. These results provide a benchmarked multi-platform computational framework for characterizing structural differences between transcription factor isoforms and establish a performance baseline for future approaches that incorporate experimental binding data into model training.| File | Dimensione | Formato | |
|---|---|---|---|
|
sara_soleimanijevinani.pdf
accesso aperto
Dimensione
2.04 MB
Formato
Adobe PDF
|
2.04 MB | Adobe PDF | Visualizza/Apri |
È consentito all'utente scaricare e condividere i documenti disponibili a testo pieno in UNITESI UNIPV nel rispetto della licenza Creative Commons del tipo CC BY NC ND.
Per maggiori informazioni e per verifiche sull'eventuale disponibilità del file scrivere a: [email protected].
https://hdl.handle.net/20.500.14239/36025