hh.sePublications
Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Group-Sparse Manifold-Aware Integrated Gradients for Multimodal Transformers on EHR Trajectories
Halmstad University, School of Information Technology.ORCID iD: 0000-0002-1999-8435
Halmstad University, School of Information Technology. Region Halland, Halmstad, Sweden.ORCID iD: 0000-0003-2006-6229
Halmstad University, School of Information Technology. Lund University, Lund, Sweden.ORCID iD: 0000-0003-1145-4297
2025 (English)In: Proceedings of Machine Learning Research, Cambridge, MA: JMLR , 2025, Vol. 297, p. 1-19Conference paper, Published paper (Refereed)
Abstract [en]

Integrated Gradients (IG) is a popular method for explaining clinical deep models—including widely used multimodal, pretrained Transformers—but its utility on EHR code sequences is hampered by (i) the lack of principled baselines for sequence of discrete tokens and (ii) dense, hard-to-interpret generated attributions. To address both, first, we introduce a manifold-aware baseline: the expected value under the empirical dist—implemented as the position-wise empirical mean of pre-Transformer token embeddings on held-out validation data, which keeps IG interpolants near the data manifold. Second, we introduce {GS-IG}, which preserves the straight path geometry but re-parameterizes the schedule (\alpha(t)=t^{\theta}) and selects (\theta) per input by minimizing a token-level (\ell_{2,1}) (group-sparsity) objective, producing concise, practitioner-friendly explanations. On MIMIC-IV (incident heart failure) and MDC (early mortality), the manifold-aware baseline improves faithfulness (higher Comprehensiveness, lower Sufficiency), and GS-IG reduces token-level (\ell_{2,1}) by 9–18% with negligible change in those metrics on the manifold-aware baseline. The method is lightweight and yields faithful, sparse, and actionable. © 2025 A. Amirahmadi, F. Etminani & M. Ohlsson.

Place, publisher, year, edition, pages
Cambridge, MA: JMLR , 2025. Vol. 297, p. 1-19
Series
Proceedings of Machine Learning Research, ISSN 2640-3498
Keywords [en]
Integrated Gradients, Explainability, Multimodal Transformers, Group Sparsity, Manifold-aware, Electronic Health Records (EHR), Patient trajectories
National Category
Computer and Information Sciences
Identifiers
URN: urn:nbn:se:hh:diva-58437OAI: oai:DiVA.org:hh-58437DiVA, id: diva2:2038832
Conference
Machine Learning for Health (ML4H) 2025, San Diego, USA, 1-2 december, 2025
Funder
Swedish Research Council, 019-00198Knowledge Foundation, 20200208 01 HAvailable from: 2026-02-16 Created: 2026-02-16 Last updated: 2026-02-19Bibliographically approved
In thesis
1. Learning More from Less: Accurate and Trustworthy Foundation Models for Patient Trajectories
Open this publication in new window or tab >>Learning More from Less: Accurate and Trustworthy Foundation Models for Patient Trajectories
2026 (English)Doctoral thesis, comprehensive summary (Other academic)
Abstract [en]

Electronic health records (EHRs) contain longitudinal traces of patients’ interactions with the healthcare system. These patient trajectories—sequences of diagnoses, medications, and other events over time—offer opportunities to predict adverse outcomes early to intervene. In practice, however, EHR data are heterogeneous, temporally complex, and often available only in limited-sized cohorts with scarce labels. This thesis, Learning More from Less: Accurate and Trustworthy Foundation Models for Patient Trajectories, investigates how to build foundation-style models for such data.

The work is guided by the question: How can we improve prediction and provide trustworthy explanations for adverse health outcomes by modeling longitudinal EHR trajectories? It follows two tracks: (i) robust EHR-specific representation learning, and (ii) trustworthy modeling. 

First, the thesis enriches self-supervised pretraining for structured EHR. A trajectory-order objective (TOO-BERT) teaches models to distinguish true temporal order from plausible permutations, while a source-masked objective model cross-sources dependencies. These objectives exploit the structure already present in trajectories, yielding stronger representations and improved prediction of incident outcomes.

Second, the thesis targets robust adaptation under label scarcity. Adaptive Noise-Augmented Attention (ANAA) perturbs and smoothly augments attention scores during fine-tuning, broadening overly sharp attention patterns and improving performance.

Third, the thesis develops explanation methods tailored to multimodal transformers EHR telemetry models. A manifold-aware baseline for Integrated Gradients keeps attribution paths in high-density regions of the representation space, improving faithfulness. Group-Sparse IG further adjusts the path schedule to produce sparse, token-level explanations that are more concise. Building on these methods, the thesis also proposes an approach to aggregate individual-level attributions into population-level insights for greater actionability, and applies it to identify key drivers of longevity and early mortality in the Malmö Diet and Cancer cohort

Finally, the thesis explores uncertainty estimation in small, sequence-based datasets through a Gaussian process model with a decoupled global alignment kernel for peptide permeability prediction. This demonstrates how structured sequence kernels can provide better accuracy and calibrated uncertainty when data are limited.

Overall, the thesis shows that in complex, data-scarce EHR settings, ``learning more from less'' requires making the pretraining, fine-tuning, and explanation stages explicitly reflect the structure of patient trajectories, leading to more accurate and trustworthy models for clinical risk prediction.

Place, publisher, year, edition, pages
Halmstad: Halmstad University Press, 2026. p. 46
Series
Halmstad University Dissertations ; 141
National Category
Electrical Engineering, Electronic Engineering, Information Engineering
Identifiers
urn:nbn:se:hh:diva-58472 (URN)978-91-90123-03-4 (ISBN)978-91-90123-04-1 (ISBN)
Public defence
2026-03-26, S3030, Kristian IV:s väg 3, Halmstad, 13:00 (English)
Opponent
Supervisors
Funder
Swedish Research Council, 2019-00198Knowledge Foundation, 20200208 01 H
Available from: 2026-02-26 Created: 2026-02-19 Last updated: 2026-02-26Bibliographically approved

Open Access in DiVA

fulltext(1757 kB)89 downloads
File information
File name FULLTEXT01.pdfFile size 1757 kBChecksum SHA-512
1a711ab494eb087fd10fd9939987a02e45d13331740ec9878c31321bdab25e86862523290837564fedc9baa85b1495c7f3dceb60d632376508d2487fa97f3890
Type fulltextMimetype application/pdf

Authority records

Amirahmadi, AliEtminani, FarzanehOhlsson, Mattias

Search in DiVA

By author/editor
Amirahmadi, AliEtminani, FarzanehOhlsson, Mattias
By organisation
School of Information Technology
Computer and Information Sciences

Search outside of DiVA

GoogleGoogle Scholar
The number of downloads is the sum of all downloads of full texts. It may include eg previous versions that are now no longer available

urn-nbn

Altmetric score

urn-nbn
Total: 8293 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf