hh.sePublications
Change search
Link to record
Permanent link

Direct link
Publications (10 of 32) Show all publications
Gustafsson, O., Lundström, J., Ohlsson, M., Stenhamre, H., Tsang, D., Pavia, J. & Ahlberg, E. (2026). Cohort profile: The Dutch wound monitor cohort and the Swedish Region Halland Integrated Platform (RHIP) wound cohort. PLOS ONE, 21(1), 1-16, Article ID e0339260..
Open this publication in new window or tab >>Cohort profile: The Dutch wound monitor cohort and the Swedish Region Halland Integrated Platform (RHIP) wound cohort
Show others...
2026 (English)In: PLOS ONE, E-ISSN 1932-6203, Vol. 21, no 1, p. 1-16, article id e0339260.Article in journal (Refereed) Published
Abstract [en]

Hard-to-heal wounds are a growing human and financial concern, constituting approximately 1–3% of the healthcare budget. Wound care is not a medical specialty and is often not prioritized within healthcare. A large portion of the cost and suffering caused by wounds has the potential to be mediated through improved knowledge and effectivised workflows. One potential way to achieve this is through the implementation of AI-tools to support clinicians in planning and executing wound care. Information-driven care is a framework for implementing AI-technology in healthcare. Wound Monitor is a Dutch database containing data collected from home-care visits conducted by wound specialists during 2005 to 2022, mostly in Limburg. It contains data of more than 17000 patients. Region Halland, Sweden, created a platform of integrated clinical, financial and operational data called “The Regional Healthcare Information Platform” (RHIP). The platform contains data on over 500 000 patients during 2008–2021. Within this data, a subset of almost 39000 patients have been diagnosed with wounds or wound related conditions. This subset of patients are defined as the RHIP Wound Cohort. This article characterizes the two wound cohorts in terms of demographics and wound types. Further, it examines the quality, quantity and granularity of the respective databases. The discussion section evaluates the strengths and weaknesses of the datasets in terms of the perspective they provide on the patient and wound journey. Lastly, the discussion section also explores how the cohorts may be utilized for predictive modeling and other machine learning-based applications in order to enable information-driven wound care. © 2026 Gustafsson et al. This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.

Place, publisher, year, edition, pages
San Francisco: Public Library of Science (PLoS), 2026
National Category
Nursing
Research subject
Health Innovation, IDC
Identifiers
urn:nbn:se:hh:diva-58282 (URN)10.1371/journal.pone.0339260 (DOI)001697761300002 ()41564038 (PubMedID)2-s2.0-105028227665 (Scopus ID)
Note

This research is included in the CAISR Health research profile.

Available from: 2026-06-11 Created: 2026-06-11 Last updated: 2026-07-06Bibliographically approved
Vinge, R., Byttner, S. & Lundström, J. (2025). Expanding Polynomial Kernels for Global and Local Explanations of Support Vector Machines. In: Advances in Intelligent Data Analysis XXIII (IDA 2025): Proceedings. Paper presented at 23rd International Symposium on Intelligent Data Analysis, IDA 2025, Konstanz, Germany, May 7–9, 2025, (pp. 456-468). Cham: Springer
Open this publication in new window or tab >>Expanding Polynomial Kernels for Global and Local Explanations of Support Vector Machines
2025 (English)In: Advances in Intelligent Data Analysis XXIII (IDA 2025): Proceedings, Cham: Springer, 2025, p. 456-468Conference paper, Published paper (Refereed)
Abstract [en]

Researchers and practitioners of machine learning nowadays rarely overlook the potential of using explainable AI methods to understand models and their predictions. These explainable AI methods mainly focus on the importance of individual input features. However, as important as the input features themselves, are the interactions between them. Methods such as the model-agnostic but computationally expensive Friedman’s H-statistic and SHAP investigate and estimate the impact of interactions between the features. Due to computational constraints, the investigation is often limited to second-order interactions. In this paper, we present a novel, model-specific method to explain the impact of feature interactions in SVM classifiers with polynomial kernels. The method is computationally frugal and calculates the interaction importance exactly for any order of interaction. Explainability is achieved by mathematical transformation to a linear model with full fidelity to the original model. Further, we show how the model provides for both global and local explanations, and facilitates post-hoc feature selection. We demonstrate the method on two datasets; one is an artificial dataset where H-statistics requires extra care to provide useful interpretation; and one on the real-world scenario of the Wisconsin Breast Cancer dataset. Our experiments show that the method provides reasonable, easy to interpret and fast to compute explanations of the trained model. © The Author(s), under exclusive license to Springer Nature Switzerland AG 2025.

Place, publisher, year, edition, pages
Cham: Springer, 2025
Series
Lecture Notes in Computer Science ; 15669
Keywords
Explainability, SVM, Trustworthy AI
National Category
Artificial Intelligence Computer Sciences Computer graphics and computer vision
Research subject
Health Innovation; Health Innovation, IDC
Identifiers
urn:nbn:se:hh:diva-56292 (URN)10.1007/978-3-031-91398-3_34 (DOI)2-s2.0-105005271620 (Scopus ID)978-3-031-91397-6 (ISBN)978-3-031-91398-3 (ISBN)
Conference
23rd International Symposium on Intelligent Data Analysis, IDA 2025, Konstanz, Germany, May 7–9, 2025,
Note

This research is included in the CAISR Health research profile.

Available from: 2025-07-08 Created: 2025-07-08 Last updated: 2025-10-01Bibliographically approved
Liang, G., Abiri, N., Hashemi, A. S., Lundström, J., Byttner, S. & Tiwari, P. (2025). Latent Space Score-based Diffusion Model for Probabilistic Multivariate Time Series Imputation. In: 2025 IEEE International Conference On Acoustics, Speech And Signal Processing (Icassp): Conference Proceedings. Paper presented at 2025 IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP 2025), April 6-11, 2025, Hyderabad, India (pp. 1-5). Piscataway: IEEE
Open this publication in new window or tab >>Latent Space Score-based Diffusion Model for Probabilistic Multivariate Time Series Imputation
Show others...
2025 (English)In: 2025 IEEE International Conference On Acoustics, Speech And Signal Processing (Icassp): Conference Proceedings, Piscataway: IEEE, 2025, p. 1-5Conference paper, Published paper (Refereed)
Abstract [en]

Accurate imputation is essential for the reliability and success of downstream tasks. Recently, diffusion models have attracted great attention in this field. However, these models neglect the latent distribution in a lower-dimensional space derived from the observed data, which limits the generative capacity of the diffusion model. Additionally, dealing with the original missing data without labels becomes particularly problematic. To address these issues, we propose the Latent Space Score-Based Diffusion Model (LSSDM) for probabilistic multivariate time series imputation. Observed values are projected onto low-dimensional latent space and coarse values of the missing data are reconstructed without knowing their ground truth values by this unsupervised learning approach. Finally, the reconstructed values are fed into a conditional diffusion model to obtain the precise imputed values of the time series. In this way, LSSDM not only possesses the power to identify the latent distribution but also seamlessly integrates the diffusion model to obtain the high-fidelity imputed values and assess the uncertainty of the dataset. Experimental results demonstrate that LSSDM achieves superior imputation performance while also providing a better explanation and uncertainty analysis of the imputation mechanism. The website of the code is https://github.com/gorgen2020/LSSDMimputation. © 2025 IEEE.

Place, publisher, year, edition, pages
Piscataway: IEEE, 2025
Series
Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing, ISSN 1520-6149, E-ISSN 2379-190X
Keywords
Diffusion model, multivariate time series, imputation, variational graph autoencoder
National Category
Probability Theory and Statistics Computer and Information Sciences
Identifiers
urn:nbn:se:hh:diva-58172 (URN)10.1109/ICASSP49660.2025.10888912 (DOI)001611514300514 ()2-s2.0-105003891405 (Scopus ID)979-8-3503-6875-8 (ISBN)979-8-3503-6874-1 (ISBN)
Conference
2025 IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP 2025), April 6-11, 2025, Hyderabad, India
Available from: 2026-01-16 Created: 2026-01-16 Last updated: 2026-01-16Bibliographically approved
Kanwal, S., Nowaczyk, S., Rahat, M., Lundström, J. & Khan, F. (2024). Deep Learning for Generating Synthetic Traffic Data. In: Xin-She Yang; Simon Sherratt; Nilanjan Dey; Amit Joshi (Ed.), Proceedings of Ninth International Congress on Information and Communication Technology: ICICT 2024, London, Volume 8. Paper presented at 9th International Congress on Information and Communication Technology, ICICT 2024, London, United Kingdom, 19-22 February, 2024 (pp. 431-454). Singapore: Springer, 1004 LNNS
Open this publication in new window or tab >>Deep Learning for Generating Synthetic Traffic Data
Show others...
2024 (English)In: Proceedings of Ninth International Congress on Information and Communication Technology: ICICT 2024, London, Volume 8 / [ed] Xin-She Yang; Simon Sherratt; Nilanjan Dey; Amit Joshi, Singapore: Springer, 2024, Vol. 1004 LNNS, p. 431-454Conference paper, Published paper (Refereed)
Abstract [en]

The purpose of the study is to demonstrate the feasibility of combining traffic simulator technology with machine learning (ML) methods to create realistic and comprehensive synthetic traffic data. Synthetic data alleviates many ethical and privacy concerns, significantly reduces the costs associated with data collection, and enables researchers to study scenarios and conditions that are difficult or impossible to replicate in real-world environments. Access to large amounts of diverse and controlled data is essential for developing and testing artificial intelligence (AI) models and leads to more reliable and robust results. Traffic simulators like SUMO have been successfully used for that purpose in the past, creating realistic vehicular traces. One drawback is that, without coupling them with complex physics emulators, they are not capable of generating internal vehicle parameters. Such parameters, on the other hand, are crucial for many purposes, from understanding energy efficiency and optimizing driver behavior to predictive maintenance and monitoring the degradation of key components, such as driveline batteries. In this paper, we propose Synthetic Traffic Data Generator (STDG) and demonstrate that an ML model that is trained on the internal parameters of a vehicle in one set of conditions (Sweden) can be used to generate synthetic data corresponding to another setting (Monaco). The proposed method promises to eliminate the need for an expensive collection of the original vehicle parameters across many different settings. Moreover, sharing the synthetic data with additional stakeholders is easier due to the reduced security and integrity risk of exposing the vehicle’s privacy-sensitive original parameters. This study compares several ML techniques, including deep learning (DL) based, for generating internal parameters of vehicles, such as fuel rate, engine speed, and wet tank air pressure. Using the actual bus data from a small city to train our ML models, we attempt to forecast the internal parameters of the buses in various scenarios. The proposed method first utilizes SUMO to generate synthetic waypoints for the bus and then predicts the other parameters using the trained model, thereby producing synthetic data with internal parameters for buses operating in a new urban environment. Our preliminary results indicated that our model is performing well within a 90% confidence interval. © The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2024.

Place, publisher, year, edition, pages
Singapore: Springer, 2024
Series
Lecture Notes in Networks and Systems, ISSN 2367-3370, E-ISSN 2367-3389 ; 1004
Keywords
Deep learning, Machine learning, Synthetic data, Traffic simulation
National Category
Computer Sciences
Identifiers
urn:nbn:se:hh:diva-54494 (URN)10.1007/978-981-97-3305-7_36 (DOI)001327002400036 ()2-s2.0-85201095610 (Scopus ID)
Conference
9th International Congress on Information and Communication Technology, ICICT 2024, London, United Kingdom, 19-22 February, 2024
Funder
Knowledge FoundationVinnova
Available from: 2024-08-26 Created: 2024-08-26 Last updated: 2025-10-01Bibliographically approved
Khoshkangini, R., Tajgardan, M., Lundström, J., Rabbani, M. & Tegnered, D. (2023). A Snapshot-Stacked Ensemble and Optimization Approach for Vehicle Breakdown Prediction. Sensors, 23(12), Article ID 5621.
Open this publication in new window or tab >>A Snapshot-Stacked Ensemble and Optimization Approach for Vehicle Breakdown Prediction
Show others...
2023 (English)In: Sensors, E-ISSN 1424-8220, Vol. 23, no 12, article id 5621Article in journal (Refereed) Published
Abstract [en]

Predicting breakdowns is becoming one of the main goals for vehicle manufacturers so as to better allocate resources, and to reduce costs and safety issues. At the core of the utilization of vehicle sensors is the fact that early detection of anomalies facilitates the prediction of potential breakdown issues, which, if otherwise undetected, could lead to breakdowns and warranty claims. However, the making of such predictions is too complex a challenge to solve using simple predictive models. The strength of heuristic optimization techniques in solving np-hard problems, and the recent success of ensemble approaches to various modeling problems, motivated us to investigate a hybrid optimization- and ensemble-based approach to tackle the complex task. In this study, we propose a snapshot-stacked ensemble deep neural network (SSED) approach to predict vehicle claims (in this study, we refer to a claim as being a breakdown or a fault) by considering vehicle operational life records. The approach includes three main modules: Data pre-processing, Dimensionality Reduction, and Ensemble Learning. The first module is developed to run a set of practices to integrate various sources of data, extract hidden information and segment the data into different time windows. In the second module, the most informative measurements to represent vehicle usage are selected through an adapted heuristic optimization approach. Finally, in the last module, the ensemble machine learning approach utilizes the selected measurements to map the vehicle usage to the breakdowns for the prediction. The proposed approach integrates, and uses, the following two sources of data, collected from thousands of heavy-duty trucks: Logged Vehicle Data (LVD) and Warranty Claim Data (WCD). The experimental results confirm the proposed system’s effectiveness in predicting vehicle breakdowns. By adapting the optimization and snapshot-stacked ensemble deep networks, we demonstrate how sensor data, in the form of vehicle usage history, contributes to claim predictions. The experimental evaluation of the system on other application domains also indicated the generality of the proposed approach. © 2023 by the authors.

Place, publisher, year, edition, pages
Basel: MDPI, 2023
Keywords
breakdown prediction, deep neural networks, ensemble learning, optimization
National Category
Mechanical Engineering
Identifiers
urn:nbn:se:hh:diva-51309 (URN)10.3390/s23125621 (DOI)001015804000001 ()37420787 (PubMedID)2-s2.0-85163933766 (Scopus ID)
Available from: 2023-08-09 Created: 2023-08-09 Last updated: 2025-10-01Bibliographically approved
Hashemi, A. S., Soliman, A., Lundström, J. & Etminani, K. (2023). Domain Knowledge-Driven Generation of Synthetic Healthcare Data. In: Maria Hägglund; Madeleine Blusi; Stefano Bonacina; Lina Nilsson; Inge Cort Madsen; Sylvia Pelayo; Anne Moen; Arriel Benis; Lars Lindsköld; Parisis Gallos (Ed.), Caring is Sharing – Exploiting the Value in Data for Health and Innovation: Proceedings of MIE 2023. Paper presented at The 33rd Medical Informatics Europe Conference, MIE2023, Gothenburg, Sweden, 22-25 May, 2023 (pp. 352-353). Amsterdam: IOS Press, 302
Open this publication in new window or tab >>Domain Knowledge-Driven Generation of Synthetic Healthcare Data
2023 (English)In: Caring is Sharing – Exploiting the Value in Data for Health and Innovation: Proceedings of MIE 2023 / [ed] Maria Hägglund; Madeleine Blusi; Stefano Bonacina; Lina Nilsson; Inge Cort Madsen; Sylvia Pelayo; Anne Moen; Arriel Benis; Lars Lindsköld; Parisis Gallos, Amsterdam: IOS Press, 2023, Vol. 302, p. 352-353Conference paper, Published paper (Refereed)
Abstract [en]

Healthcare longitudinal data collected around patients' life cycles, today offer a multitude of opportunities for healthcare transformation utilizing artificial intelligence algorithms. However, access to "real" healthcare data is a big challenge due to ethical and legal reasons. There is also a need to deal with challenges around electronic health records (EHRs) including biased, heterogeneity, imbalanced data, and small sample sizes. In this study, we introduce a domain knowledge-driven framework for generating synthetic EHRs, as an alternative to methods only using EHR data or expert knowledge. By leveraging external medical knowledge sources in the training algorithm, the suggested framework is designed to maintain data utility, fidelity, and clinical validity while preserving patient privacy. © 2023 European Federation for Medical Informatics (EFMI) and IOS Press.

Place, publisher, year, edition, pages
Amsterdam: IOS Press, 2023
Series
Studies in Health Technology and Informatics, ISSN 0926-9630, E-ISSN 1879-8365 ; 302
Keywords
Domain Knowledge, EHR, Representation Learning, Synthetic Data
National Category
Computer Sciences
Research subject
Health Innovation, IDC
Identifiers
urn:nbn:se:hh:diva-51733 (URN)10.3233/SHTI230136 (DOI)37203680 (PubMedID)2-s2.0-85159760846 (Scopus ID)978-1-64368-389-8 (ISBN)
Conference
The 33rd Medical Informatics Europe Conference, MIE2023, Gothenburg, Sweden, 22-25 May, 2023
Note

This research is included in the CAISR Health research profile.

Available from: 2023-10-03 Created: 2023-10-03 Last updated: 2025-10-01Bibliographically approved
Lundström, J., Hashemi, A. S. & Tiwari, P. (2023). Explainable Graph Neural Networks for Atherosclerotic Cardiovascular Disease. In: Caring is sharing - exploiting the value in data for health and innovation: [33rd Medical Informatics Europe Conference, MIE2023, held in Gothenburg, Sweden, from 22 to 25 May. Paper presented at 33rd Medical Informatics Europe Conference: Caring is Sharing - Exploiting the Value in Data for Health and Innovation, MIE2023, Gothenburg, 22-25 May 2023, Code 189285 (pp. 603-604). Amsterdam: IOS Press, 302
Open this publication in new window or tab >>Explainable Graph Neural Networks for Atherosclerotic Cardiovascular Disease
2023 (English)In: Caring is sharing - exploiting the value in data for health and innovation: [33rd Medical Informatics Europe Conference, MIE2023, held in Gothenburg, Sweden, from 22 to 25 May, Amsterdam: IOS Press, 2023, Vol. 302, p. 603-604Conference paper, Published paper (Refereed)
Abstract [en]

Understanding the aspects of progression for atherosclerotic cardiovascular disease and treatment is key to building reliable clinical decision-support systems. To promote system trust, one step is to make the machine learning models (used by the decision support systems) explainable for clinicians, developers, and researchers. Recently, working with longitudinal clinical trajectories using Graph Neural Networks (GNNs) has attracted attention among machine learning researchers. Although GNNs are seen as black-box methods, promising explainable AI (XAI) methods for GNNs have lately been proposed. In this paper, which describes initial project stages, we aim at utilizing GNNs for modeling, predicting, and exploring the model explainability of the low-density lipoprotein cholesterol level in long-term atherosclerotic cardiovascular disease progression and treatment.

Place, publisher, year, edition, pages
Amsterdam: IOS Press, 2023
Series
Studies in Health Technology and Informatics, ISSN 1879-8365, E-ISSN 1879-8365 ; 302
Keywords
Cardiovascular Diseases, EHR, Graph Neural Networks
National Category
Neurosciences
Identifiers
urn:nbn:se:hh:diva-51975 (URN)10.3233/SHTI230214 (DOI)001071432900157 ()37203757 (PubMedID)2-s2.0-85159762049 (Scopus ID)9781643683881 (ISBN)
Conference
33rd Medical Informatics Europe Conference: Caring is Sharing - Exploiting the Value in Data for Health and Innovation, MIE2023, Gothenburg, 22-25 May 2023, Code 189285
Note

This research is included in the CAISR Health research profile.

Available from: 2023-11-13 Created: 2023-11-13 Last updated: 2025-10-01Bibliographically approved
Khoshkangini, R., Sheikholharam Mashhadi, P., Tegnered, D., Lundström, J. & Rögnvaldsson, T. (2023). Predicting Vehicle Behavior Using Multi-task Ensemble Learning. Expert systems with applications, 212, Article ID 118716.
Open this publication in new window or tab >>Predicting Vehicle Behavior Using Multi-task Ensemble Learning
Show others...
2023 (English)In: Expert systems with applications, ISSN 0957-4174, E-ISSN 1873-6793, Vol. 212, article id 118716Article in journal (Refereed) Published
Abstract [en]

Vehicle utilization analysis is an essential tool for manufacturers to understand customer needs, improve equipment uptime, and to collect information for future vehicle and service development. Typically today, this behavioral modeling is done on high-resolution time-resolved data with features such as GPS position and fuel consumption. However, high-resolution data is costly to transfer and sensitive from a privacy perspective. Therefore, such data is typically only collected when the customer pays for extra services relying on that data. This motivated us to develop a multi-task ensemble approach to transfer knowledge from the high-resolution data and enable vehicle behavior prediction from low-resolution but high dimensional data that is aggregated over time in the vehicles.

This study proposes a multi-task snapshot-stacked ensemble (MTSSE) deep neural network for vehicle behavior prediction by considering vehicles’ low-resolution operational life records. The multi-task ensemble approach utilizes the measurements to map the low-frequency vehicle usage to the vehicle behaviors defined from the high-resolution time-resolved data. Two data sources are integrated and used: high-resolution data called Dynafleet, and low-resolution so-called Logged Vehicle Data (LVD). The experimental results demonstrate the proposed approach’s effectiveness in predicting the vehicle behavior from low frequency data. With the suggested multi-task snapshot-stacked ensemble deep network, it is shown how low-resolution sensor data can highly contribute to predicting multiple vehicle behaviors simultaneously while using only one single training process. © 2022 The Author(s)

Place, publisher, year, edition, pages
Oxford: Elsevier, 2023
Keywords
Behavior modeling, Multi-task learning, Deep neural networks, Ensemble learning
National Category
Information Systems
Identifiers
urn:nbn:se:hh:diva-48170 (URN)10.1016/j.eswa.2022.118716 (DOI)000870841300003 ()2-s2.0-85138456634 (Scopus ID)
Available from: 2022-09-29 Created: 2022-09-29 Last updated: 2025-10-01Bibliographically approved
Hashemi, A. S., Etminani, K., Soliman, A., Hamed, O. & Lundström, J. (2023). Time-series Anonymization of Tabular Health Data using Generative Adversarial Network. In: 2023 International Joint Conference on Neural Networks (IJCNN): . Paper presented at 2023 International Joint Conference on Neural Networks, IJCNN 2023, Gold Coast, Queensland, Australia, 18-23 June, 2023. Piscataway, NJ: IEEE
Open this publication in new window or tab >>Time-series Anonymization of Tabular Health Data using Generative Adversarial Network
Show others...
2023 (English)In: 2023 International Joint Conference on Neural Networks (IJCNN), Piscataway, NJ: IEEE, 2023Conference paper, Published paper (Refereed)
Abstract [en]

Data anonymization has been used as a fundamental tool in various domains, e.g. healthcare, to alter personal data such that individuals can no longer be identified directly or indirectly in a way to enable broader sharing of data. For example, data perturbation techniques add noise to original data allowing individual record confidentiality while maintaining high-quality data for analytical purposes. In this paper, we propose a perturbation technique for anonymizing longitudinal tabular data such as electronic health records (EHRs). Our model starts by learning a latent space of original data to better capture temporal trends, then employs a generative adversarial network together to train a perturbation generator. During model training, a time-supervised loss function for handling sequence-dependent noise, together with the adversarial unsupervised, anonymization, and reconstruction loss functions are utilized. To evaluate our model quantitatively, we use multiple evaluation metrics for the fidelity, utility, and identifiability of generated data, in addition, the model is evaluated qualitatively by visualizing generated and original data. The results confirm that our model preserves the privacy of the original data and generates a perturbed version with high fidelity and utility compared to some state-of-the-art techniques. © 2023 IEEE.

Place, publisher, year, edition, pages
Piscataway, NJ: IEEE, 2023
Series
Proceedings of ... International Joint Conference on Neural Networks, ISSN 2161-4393, E-ISSN 2161-4407
Keywords
anonymization, data perturbation, EHR, generative adversarial networks, synthetic data
National Category
Computer Sciences
Research subject
Health Innovation, IDC
Identifiers
urn:nbn:se:hh:diva-51673 (URN)10.1109/IJCNN54540.2023.10191367 (DOI)2-s2.0-85169602590 (Scopus ID)9781665488679 (ISBN)9781665488686 (ISBN)
Conference
2023 International Joint Conference on Neural Networks, IJCNN 2023, Gold Coast, Queensland, Australia, 18-23 June, 2023
Note

This research is included in the CAISR Health research profile.

Available from: 2023-09-22 Created: 2023-09-22 Last updated: 2025-10-01Bibliographically approved
Ali Hamad, R., Kimura, M. & Lundström, J. (2020). Efficacy of Imbalanced Data Handling Methods on Deep Learning for Smart Homes Environments. SN Computer Science, 1(4), Article ID 204.
Open this publication in new window or tab >>Efficacy of Imbalanced Data Handling Methods on Deep Learning for Smart Homes Environments
2020 (English)In: SN Computer Science, E-ISSN 2661-8907, Vol. 1, no 4, article id 204Article in journal (Refereed) Published
Abstract [en]

Human activity recognition as an engineering tool as well as an active research field has become fundamental to many applications in various fields such as health care, smart home monitoring and surveillance. However, delivering sufficiently robust activity recognition systems from sensor data recorded in a smart home setting is a challenging task. Moreover, human activity datasets are typically highly imbalanced because generally certain activities occur more frequently than others. Consequently, it is challenging to train classifiers from imbalanced human activity datasets. Deep learning algorithms perform well on balanced datasets, yet their performance cannot be promised on imbalanced datasets. Therefore, we aim to address the problem of class imbalance in deep learning for smart home data. We assess it with Activities of Daily Living recognition using binary sensors dataset. This paper proposes a data level perspective combined with a temporal window technique to handle imbalanced human activities from smart homes in order to make the learning algorithms more sensitive to the minority class. The experimental results indicate that handling imbalanced human activities from the data-level outperforms algorithms level and improved the classification performance. © The Author(s) 2020

Place, publisher, year, edition, pages
Heidelberg: Springer Berlin/Heidelberg, 2020
Keywords
Activity recognition, Smart home, Imbalanced class
National Category
Computer Sciences
Identifiers
urn:nbn:se:hh:diva-42533 (URN)10.1007/s42979-020-00211-1 (DOI)2-s2.0-85089693380 (Scopus ID)
Funder
Knowledge Foundation, 20100271
Note

Funding: Open access funding provided by Halmstad University. This research is supported by the Knowledge Foundation under the project of the Center for Applied Intelligent Systems, under Grant Agreement No. 20100271.

Available from: 2020-06-19 Created: 2020-06-19 Last updated: 2026-06-05Bibliographically approved
Organisations
Identifiers
ORCID iD: ORCID iD iconorcid.org/0000-0001-8804-5884

Search in DiVA

Show all publications