Digitala Vetenskapliga Arkivet

Endre søk
RefereraExporteraLink to record
Permanent link

Direct link
Referera
Referensformat
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Annet format
Fler format
Språk
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Annet språk
Fler språk
Utmatningsformat
  • html
  • text
  • asciidoc
  • rtf
Domain Generalization of Deep Learning Models Under Subgroup Shift in Breast Cancer Prognosis
Högskolan i Skövde, Institutionen för biovetenskap. Högskolan i Skövde, Forskningsmiljön Systembiologi. (Translational Bioinformatics)ORCID-id: 0000-0003-4191-8435
Högskolan i Skövde, Institutionen för biovetenskap. Högskolan i Skövde, Forskningsmiljön Systembiologi. (Translational Bioinformatics)ORCID-id: 0000-0001-9242-4852
Högskolan i Skövde, Institutionen för informationsteknologi. Högskolan i Skövde, Forskningsmiljön Informationsteknologi. (Skövde Artificial Intelligence Lab (SAIL))ORCID-id: 0000-0001-8884-2154
Högskolan i Skövde, Institutionen för biovetenskap. Högskolan i Skövde, Forskningsmiljön Systembiologi. Institute of Medicine, Sahlgrenska Academy University of Gothenburg, Sweden. (Translational Bioinformatics)ORCID-id: 0000-0003-4697-0590
2024 (engelsk)Inngår i: 2024 IEEE Conference on Computational Intelligence in Bioinformatics and Computational Biology (CIBCB), IEEE, 2024Konferansepaper, Publicerat paper (Fagfellevurdert)
Abstract [en]

Making breast cancer prognosis from gene expression profiles of the primary tumor has become a promising application of deep learning. Yet, to be relevant to real world applications in the clinic and for knowledge discovery, these models must be robust to common distribution shifts. In this study, we evaluate recently proposed methods for improving domain and subgroup shifts. We test the in-distribution and out-of-distribution generalization of multiple episode learning, stochastic weight averaging, group distributionally robust optimization, and a subsampling scheme on one training and four external breast cancer prognosis datasets. The evaluation found that the methods can, to various degrees, improve generalization across domains, although there remain, partially high, generalization gaps. Additionally, in-distribution and out-of-distribution generalization differs between clinical subtypes of breast cancer. Thus, we conclude that further research into methods specifically addressing challenges in breast cancer prognosis from gene expression data are warranted. 

sted, utgiver, år, opplag, sider
IEEE, 2024.
Serie
IEEE Conference on Computational Intelligence in Bioinformatics and Computational Biology (CIBCB), ISSN 2994-9351, E-ISSN 2994-9408
Emneord [en]
breast cancer, domain generalization, gene expression, subgroup shift, survival analysis, Contrastive Learning, Diseases, Lung cancer, Stochastic systems, Breast cancer prognosis, Gene expression profiles, Generalisation, Genes expression, Learning models, Real-world
HSV kategori
Forskningsprogram
Bioinformatik; Skövde Artificial Intelligence Lab (SAIL)
Identifikatorer
URN: urn:nbn:se:his:diva-24659DOI: 10.1109/CIBCB58642.2024.10702166ISI: 001546450400010Scopus ID: 2-s2.0-85207504799ISBN: 979-8-3503-5663-2 (digital)ISBN: 979-8-3503-5664-9 (tryckt)OAI: oai:DiVA.org:his-24659DiVA, id: diva2:1911226
Konferanse
21st IEEE Conference on Computational Intelligence in Bioinformatics and Computational Biology, CIBCB 2024, 27-29 August 2024, Natal, Brazil
Forskningsfinansiär
Knowledge Foundation, 20170302Knowledge Foundation, 20200014Swedish Research Council, 2022-06725
Merknad

© 2024 IEEE

Correspondence Address: S.R. Stahlschmidt; University of Skövde, Systems Biology Research Center, Skövde, Sweden; email: soren.richard.stahlschmidt@his.se

This work was supported by the University of Skövde, Sweden under grants from the Knowledge Foundation (20170302, 20200014). The computations were enabled by resources provided by Chalmers e-Commons at Chalmers and the National Academic Infrastructure for Supercomputing in Sweden (NAISS), partially funded by the Swedish Research Council through grant agreement no. 2022-06725.

Tilgjengelig fra: 2024-11-07 Laget: 2024-11-07 Sist oppdatert: 2025-10-17bibliografisk kontrollert
Inngår i avhandling
1. Machine Learning for Predicting Cancer Endpoints from Bulk Omics Data: Generalizing Knowledge from Various Modalities Across Domains
Åpne denne publikasjonen i ny fane eller vindu >>Machine Learning for Predicting Cancer Endpoints from Bulk Omics Data: Generalizing Knowledge from Various Modalities Across Domains
2025 (engelsk)Doktoravhandling, med artikler (Annet vitenskapelig)
Abstract [en]

Cancer remains one of the leading causes of death and is a major burden on patients and healthcare systems. One difficulty for finding effective treatment and matching patients to the right treatment strategy is the complexity of tumor biology. Machine learning holds the potential to learn patterns from data generated by high-throughput technologies, such as RNA-sequencing, that can elucidate the mechanisms underlying cancers and make clinically relevant predictions. In this thesis, we investigate the modeling of cancer with machine learning approaches from different molecular perspectives. First, we review the literature on the fusion of biomedical modalities with multimodal deep neural networks. In this review, we provide a descriptive overview, propose a novel taxonomy, and identify relevant research gaps. Moreover, for models to be applicable to clinical practice, they must be robust to shifts in the distribution patients are sampled from. Such shifts can stem from differences in the underlying biology or technical variation introduced during the processing of the biological material. Therefore, in two studies, we investigate domain generalization of machine learning models trained with bulk RNA-sequencing data to predict cancer survival endpoints. First, we show that deep learning-based domain generalization methods developed on non-molecular data improve robustness to distributional shifts on molecular data. We test these methods by predicting overall and recurrence free survival of breast cancer patients with subgroup shifts between source and target domains. Next, we show that relative representations of normalized count values, such as binning or ranking of expression values within a single sample, can increase domain generalization. We test these approaches in three experiments on breast, brain, and ovarian cancer. In a final study, we show that cancer stage can be predicted from circulating microRNA data with machine learning models, providing a proof of concept for this application. Overall, the work in this thesis supports making machine learning models more applicable to clinical practice by providing empirical evidence of methods improving the modeling of cancer biology. Continuing to study domain generalization of models in clinical practice and to develop methods for robustness are highlighted as future work.

sted, utgiver, år, opplag, sider
Skövde: University of Skövde, 2025. s. xi, 147
Serie
Dissertation Series ; 63
HSV kategori
Forskningsprogram
Bioinformatik
Identifikatorer
urn:nbn:se:his:diva-25131 (URN)978-91-987907-9-5 (ISBN)978-91-989080-0-8 (ISBN)
Disputas
2025-06-04, G110, University of Skövde Building G, Skövde, 13:00 (engelsk)
Opponent
Veileder
Merknad

Ett av fyra delarbeten (övriga se rubriken Delarbeten/List of papers):

3. Stahlschmidt, Sören Richard, Synnergren, Jane, and Giovannucci, Andrea (2025). “Relative Representations of RNA-seq Data Improve Domain Generalization of Machine Learning Models for Cancer Prognosis”. In: Under Submission.

Publications with low relevance:

5. Johansson, Markus, Stahlschmidt, Sören Richard, Heydarkhan-Hagvall, Sepideh, Jeppsson, Anders, Holmgren, Gustav, Sartipy, Peter, and Synnergren, Jane (2025). “Uncovering the transcriptomic landscape of cardiac hypertrophy using single-cell RNA sequencing and machine learning”. In: Under Submission.

6. Lyubetskaya, Anna et al. (2025). “In situ multi-modal characterization of pancreatic cancer reveals tumor cell identity as a defining factor of the surrounding microenvironment”. In: Under Submission.

7. Marzec-Schmidt, Katarzyna, Ghosheh, Nidal, Stahlschmidt, Sören Richard, Küppers-Munther, Barbara, Synnergren, Jane, and Ulfenborg, Benjamin (2023). “Artificial Intelligence Supports Automated Characterization of Differentiated Human Pluripotent Stem Cells”. In: Stem Cells 41.9, pp. 850–861. DOI:10. 1093/stmcls/sxad049. 

Tilgjengelig fra: 2025-05-12 Laget: 2025-05-09 Sist oppdatert: 2025-09-29bibliografisk kontrollert

Open Access i DiVA

Fulltekst mangler i DiVA

Andre lenker

Forlagets fulltekstScopus

Søk i DiVA

Av forfatter/redaktør
Stahlschmidt, Sören RichardUlfenborg, BenjaminFalkman, GöranSynnergren, Jane
Av organisasjonen

Søk utenfor DiVA

GoogleGoogle Scholar

doi
isbn
urn-nbn

Altmetric

doi
isbn
urn-nbn
Totalt: 973 treff
RefereraExporteraLink to record
Permanent link

Direct link
Referera
Referensformat
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Annet format
Fler format
Språk
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Annet språk
Fler språk
Utmatningsformat
  • html
  • text
  • asciidoc
  • rtf