hh.sePublications
Change search
Link to record
Permanent link

Direct link
Publications (10 of 148) Show all publications
Wang, M., Zhang, Y., Ren, C., Li, Q., Tiwari, P., Wang, B. & Qin, J. (2026). Adaptive Boosting LLMs for Text Classification. IEEE Transactions on Neural Networks and Learning Systems, 37(6), 2770-2781
Open this publication in new window or tab >>Adaptive Boosting LLMs for Text Classification
Show others...
2026 (English)In: IEEE Transactions on Neural Networks and Learning Systems, ISSN 2162-237X, E-ISSN 2162-2388, Vol. 37, no 6, p. 2770-2781Article in journal (Refereed) Published
Abstract [en]

With large-scale language models demonstrating superior capabilities in a wide range of downstream natural language processing tasks, the future trajectory of research in the field of text categorization faces increasing uncertainty. In this evolving paradigm of open-ended language modeling, where task delimitations are increasingly blurred, a pressing question arises: to what extent has text classification advanced under the full potential of large language model (LLM)? To address this pivotal inquiry, we introduce recurrent generative pre-trained transformer (RGPT), an adaptive boosting framework meticulously designed to craft a dedicated LLM for text classification. RGPT constructs a sequence of base learners by dynamically modulating the training data distribution and iteratively fine-tuning LLMs. These base learners are then progressively integrated, leveraging historical prediction trajectories to form a highly specialized text classification model. Extensive empirical evaluations demonstrate that RGPT surpasses eight state-of-the-art pretrained language models and seven cutting-edge LLMs across four benchmark datasets, achieving an average performance gain of 2.90%. © 2012 IEEE.

Place, publisher, year, edition, pages
Piscataway, NJ: IEEE, 2026
Keywords
Boosting, large language model, natural language processing, text classification
National Category
Natural Language Processing
Identifiers
urn:nbn:se:hh:diva-58202 (URN)10.1109/TNNLS.2025.3639613 (DOI)001663478400001 ()41525530 (PubMedID)2-s2.0-105027441383 (Scopus ID)
Note

Funding information: This work was supported in part by the National Social Science Fund of China under Grant 25AYY001; in part by the Collaborative Research with World-Leading Research Groups Scheme in The Hong Kong Polytechnic University under Project G-SACF; in part by the General Research Fund through Hong Kong Research Grants Council under Project 15218521; in part by the Natural Science Foundation of Henan Province of China under Grant 242300421412; and in part by the Foundation of Key Laboratory of Dependable Service Computing in CyberPhysical-Society (Ministry of Education), Chongqing University under Grant CPSDSC202103.

Available from: 2026-01-30 Created: 2026-01-30 Last updated: 2026-07-06Bibliographically approved
Feng, L., Li, Q., Wang, H., Li, X., Xie, F., He, F., . . . Hou, J. (2026). Advancing oral leukoplakia progression recognition: A benchmark with dataset, method, and application. Neural Networks, 195, Article ID 108256.
Open this publication in new window or tab >>Advancing oral leukoplakia progression recognition: A benchmark with dataset, method, and application
Show others...
2026 (English)In: Neural Networks, ISSN 0893-6080, E-ISSN 1879-2782, Vol. 195, article id 108256Article in journal (Refereed) Published
Abstract [en]

Oral leukoplakia, a potentially malignant disorder, is a critical precursor to oral squamous cell carcinoma (OSCC), which accounts for 90% of oral cancer cases. Early recognition of leukoplakia progression is essential for timely intervention and improved patient outcomes. However, most existing methods focus on easy oral cancer recognition, neglecting the nuanced challenge of leukoplakia progression recognition. In this paper, we introduce Oral Leukoplakia Progression Recognition (OLPR) as a novel benchmark task to classify oral lesions into three clinically relevant categories: normal, leukoplakia, and leukoplakia with cancer. We construct the OLPR dataset, a high-quality, annotated collection derived from multiple public datasets, and establish an external validation dataset using university and clinical data. In addition, we propose the Oral Leukoplakia Progression Network (OLPNet), which combines a ConvNeXt backbone pre-trained on large-scale datasets with a Feature Refinement Module (FRM) to enhance feature extraction and refinement for subtle and complex variations in oral lesion progression. Extensive experiments demonstrate the superiority of OLPNet over state-of-the-art methods, achieving 91.34% and 90.63 % F1-scores on OLPR and external datasets, respectively. Furthermore, we provide a comprehensive benchmark by evaluating over 15 classic and state-of-the-art classification models, offering valuable insights into advancing oral leukoplakia progression recognition. The dataset and related resources are released at https://github.com/qkee-lz/OLPR. © 2025 Elsevier Ltd. All rights are reserved, including those for text and data mining, AI training, and similar technologies.

Place, publisher, year, edition, pages
Oxford: Elsevier, 2026
Keywords
Oral leukoplakia progression recognition, Benchmark dataset, Methods evaluation
National Category
Odontology
Identifiers
urn:nbn:se:hh:diva-59087 (URN)10.1016/j.neunet.2025.108256 (DOI)001619399200004 ()41237693 (PubMedID)2-s2.0-105021651689 (Scopus ID)
Note

Funding information: This work is funded by the Research Fund of Anhui Institute of translational medicine (No. 2023zhyx-C46).

Available from: 2026-06-11 Created: 2026-06-11 Last updated: 2026-06-11Bibliographically approved
Wang, C., Hu, Z., Fang, X., Yu, Z. Y., Wu, Y., Xu, M., . . . Tiwari, P. (2026). Biologically-Inspired Evolutionary Domain Symbiosis for Few-shot and Zero-shot Point Cloud Semantic Segmentation. In: Sven Koenig; Chad Jenkins; Matthew E. Taylor (Ed.), Proceedings of the 40th Annual AAAI Conference on Artificial Intelligence: . Paper presented at The Fortieth AAAI Conference on Artificial Intelligence, Singapore, Singapore, 20-27 January, 2026 (pp. 9666-9674). New York: AAAI Press, 40(12)
Open this publication in new window or tab >>Biologically-Inspired Evolutionary Domain Symbiosis for Few-shot and Zero-shot Point Cloud Semantic Segmentation
Show others...
2026 (English)In: Proceedings of the 40th Annual AAAI Conference on Artificial Intelligence / [ed] Sven Koenig; Chad Jenkins; Matthew E. Taylor, New York: AAAI Press, 2026, Vol. 40, no 12, p. 9666-9674Conference paper, Published paper (Refereed)
Abstract [en]

Few-shot and zero-shot point cloud semantic segmentation aim to accurately segment novel categories using limited or no labeled samples, respectively. However, existing methods face significant challenges including domain shifts between support and query sets and the inability to handle both few-shot and zero-shot scenarios within a unified framework. To address these issues, we propose a biologically-inspired Evolutionary Domain Symbiosis Network EDS-Net for unified few-shot and zero-shot point cloud semantic segmentation. Specifically, inspired by natural symbiotic evolution, we propose a Symbiotic Evolution Module (SEM) that models co-adaptation between support and query features through self-correlation and cross-correlation mechanisms. Second, motivated by genetic crossover mechanisms, we introduce a Vision-Semantic Bridging Module (VSBM) that treats visual prototypes and semantic prototypes as two “parent” individuals, creating fused offspring prototypes through adaptive crossover operations and mutation strategies for zero-shot scenarios. Third, we develop a multi-generational evolutionary optimization framework employing an adaptive gating network to learn optimal fusion weights across different evolutionary stages. Extensive experiments demonstrate that EDS-Net with biological interpretability achieves state-of-the-art performance on both few-shot and zero-shot tasks. © 2026, Association for the Advancement of Artificial Intelligence (www.aaai.org). All rights reserved.

Place, publisher, year, edition, pages
New York: AAAI Press, 2026
Series
AAAI / IAAI proceedings, ISSN 2159-5399, E-ISSN 2374-3468 ; 12
National Category
Computer and Information Sciences
Identifiers
urn:nbn:se:hh:diva-58828 (URN)10.1609/aaai.v40i12.37929 (DOI)2-s2.0-105034566420 (Scopus ID)9781577359067 (ISBN)
Conference
The Fortieth AAAI Conference on Artificial Intelligence, Singapore, Singapore, 20-27 January, 2026
Available from: 2026-06-09 Created: 2026-06-09 Last updated: 2026-06-09Bibliographically approved
Qu, Z., Li, Y., Liu, B., Gupta, D. & Tiwari, P. (2026). DTQFL: A Digital Twin-Assisted Quantum Federated Learning Algorithm for Intelligent Diagnosis in 5G Mobile Network. IEEE journal of biomedical and health informatics, 30(1), 17-26
Open this publication in new window or tab >>DTQFL: A Digital Twin-Assisted Quantum Federated Learning Algorithm for Intelligent Diagnosis in 5G Mobile Network
Show others...
2026 (English)In: IEEE journal of biomedical and health informatics, ISSN 2168-2194, E-ISSN 2168-2208, Vol. 30, no 1, p. 17-26Article in journal (Refereed) Published
Abstract [en]

Smart healthcare aims to revolutionize med-ical services by integrating artificial intelligence (AI). The limitations of classical machine learning include privacy concerns that prevent direct data sharing among medical institutions, untimely updates, and long training times. To address these issues, this study proposes a digital twin-assisted quantum federated learning algorithm (DTQFL). By leveraging the 5G mobile network, digital twins (DT) of patients can be created instantly using data from various Internet of Medical Things (IoMT) devices and simultane-ously reduce communication time in federated learning (FL) at the same time. DTQFL generates DT for patients with specific diseases, allowing for synchronous training and updating of the variational quantum neural network (VQNN) without disrupting the VQNN in the real world. This study utilized DTQFL to train its own personalized VQNN for each hospital, considering privacy security and training speed. Simultaneously, the personalized VQNN of each hospital was obtained through further local iterations of the final global parameters. The results indicate that DTQFL can train a good VQNN without collecting local data while achieving accuracy comparable to that of data-centralized algorithms. In addition, after personalized train-ing, the VQNN can achieve higher accuracy than that with-out personalized training. © 2026 IEEE.

Place, publisher, year, edition, pages
Piscataway, NJ: Institute of Electrical and Electronics Engineers (IEEE), 2026
Keywords
digital twin, federated learning, Federated learning, Hospitals, Medical services, mobile network, Privacy, Quantum cascade lasers, quantum neural network, Servers, Smart healthcare, Training
National Category
Other Computer and Information Science
Identifiers
urn:nbn:se:hh:diva-51549 (URN)10.1109/JBHI.2023.3303401 (DOI)001662927800002 ()37552590 (PubMedID)2-s2.0-85167839904 (Scopus ID)
Available from: 2023-08-31 Created: 2023-08-31 Last updated: 2026-02-06Bibliographically approved
Sun, L., Chen, Q., Zheng, M., Ning, X., Gupta, D. & Tiwari, P. (2026). Energy-Efficient Online Continual Learning for Time Series Classification in Nanorobot-Based Smart Health. IEEE journal of biomedical and health informatics, 30(1), 81-89
Open this publication in new window or tab >>Energy-Efficient Online Continual Learning for Time Series Classification in Nanorobot-Based Smart Health
Show others...
2026 (English)In: IEEE journal of biomedical and health informatics, ISSN 2168-2194, E-ISSN 2168-2208, Vol. 30, no 1, p. 81-89Article in journal (Refereed) Published
Abstract [en]

Nanorobots have been used in smart health to collect time series data such as electrocardiograms and electroencephalograms. Real-time classification of dynamic time series signals in nanorobots is a challenging task. Nanorobots in the nanoscale range require a classification algorithm with low computational complexity. First, the classification algorithm should be able to dynamically analyze time series signals and update itself to process the concept drifts (CD). Second, the classification algorithm should have the ability to handle catastrophic forgetting (CF) and classify historical data. Most importantly, the classification algorithm should be energy-efficient to use less computing power and memory to classify signals in real-time on a smart nanorobot. To solve these challenges, we design an algorithm that can Prevent Concept Drift in Online continual Learning for time series classification (PCDOL). The prototype suppression item in PCDOL can reduce the impact caused by CD. It also solves the CF problem through the replay feature. The computation per second and the memory consumed by PCDOL are only 3.572M and 1KB, respectively. The experimental results show that PCDOL is better than several state-of-the-art methods for dealing with CD and CF in energy-efficient nanorobots. © 2026 IEEE.

Place, publisher, year, edition, pages
Piscataway, NJ: Institute of Electrical and Electronics Engineers (IEEE), 2026
Keywords
Classification algorithms, concept drift, Feature extraction, Nanobioscience, nanorobot, online continual learning, Prototypes, sensor time series classification, smart health, Task analysis, Time series analysis, Training
National Category
Computer Sciences
Identifiers
urn:nbn:se:hh:diva-51429 (URN)10.1109/JBHI.2023.3289992 (DOI)001662927800003 ()37368802 (PubMedID)2-s2.0-85163564312 (Scopus ID)
Note

Funding: National Natural Science Foundation of China (Grant Number: 61702274) and Major Key Project of PCL (Grant Number: PCL2022A03, PCL2021A02 and PCL2021A09)

Available from: 2023-08-17 Created: 2023-08-17 Last updated: 2026-02-06Bibliographically approved
Li, B., Zhang, X., Huang, Z., Tiwari, P., Zou, Q., Ding, Y. & Guo, X. (2026). Enhancing anticancer peptide discovery: A Fusion-Centric Framework With Conditional Diffusion For Prediction And Generation. PloS Computational Biology, 22(3), 1-29, Article ID e1014098.
Open this publication in new window or tab >>Enhancing anticancer peptide discovery: A Fusion-Centric Framework With Conditional Diffusion For Prediction And Generation
Show others...
2026 (English)In: PloS Computational Biology, ISSN 1553-734X, E-ISSN 1553-7358, Vol. 22, no 3, p. 1-29, article id e1014098Article in journal (Refereed) Published
Abstract [en]

Anticancer peptides (ACPs) are short bioactive sequences that selectively target tumor cells with minimal toxicity, positioning them as promising candidates for next-generation cancer therapies. However, existing computational models face limitations in sequence representation and class imbalance. To address these challenges, we propose UACD-ACPs, a unified fusion-driven framework that integrates a diffusion-inspired noise-conditioned classifier for ACP prediction and a diffusion-based peptide generation module with cancer-type-aware organization for targeted downstream screening. The classification module integrates ProtBERT-based semantic embeddings with physicochemical descriptors via the Multiscale Embedding Compression Strategy (MECS) and a diffusion-inspired noise-conditioned encoder, substantially enhancing predictive robustness and accuracy, particularly under challenging imbalanced multi-class settings. In the generative pipeline, we introduce a denoising diffusion-based generative framework augmented by two novel fusion modules: the Bitemporal Fusion Module (BFM) and the Temporal Feature Attention Module (TFAM). These modules perform multi-scale temporal and semantic fusion to promote the generation of structurally coherent and functionally relevant peptide candidates. Experimental results demonstrate that UACD-ACPs outperforms state-of-the-art methods in terms of accuracy, F1-score, and AUC-ROC. The generated peptides exhibit favorable physicochemical properties, diverse secondary structures, and strong structural stability, as validated by molecular dynamics simulations and membrane-binding analyses. Overall, this study highlights the potential of fusion-driven diffusion-based frameworks for alleviating class imbalance and data heterogeneity in anticancer peptide modeling, paving the way for scalable and biologically grounded ACP discovery. © 2026 Li et al.

Place, publisher, year, edition, pages
San Francisco: Public Library of Science (PLoS), 2026
Keywords
identification, inhibitor, language
National Category
Bioinformatics (Computational Biology) Bioinformatics (Computational Biology)
Identifiers
urn:nbn:se:hh:diva-58706 (URN)10.1371/journal.pcbi.1014098 (DOI)001724448300001 ()41886705 (PubMedID)2-s2.0-105034373461 (Scopus ID)
Available from: 2026-04-07 Created: 2026-04-07 Last updated: 2026-04-27Bibliographically approved
Jankowska, J., Kostek, B., Alonso-Fernandez, F. & Tiwari, P. (2026). Exploring the correlation between the type of music and the emotions evoked: A study using subjective questionnaires and EEG. In: Progress in Artificial Intelligence and Pattern Recognition: 9th International Congress, IWAIPR 2025, Varadero, Cuba, October 14–17, 2025, Proceedings. Paper presented at 9th International Congress, IWAIPR: International Congress on Artificial Intelligence and Pattern Recognition, October 14–17, 2025, Varadero, Cuba (pp. 395-406). Heidelberg: Springer
Open this publication in new window or tab >>Exploring the correlation between the type of music and the emotions evoked: A study using subjective questionnaires and EEG
2026 (English)In: Progress in Artificial Intelligence and Pattern Recognition: 9th International Congress, IWAIPR 2025, Varadero, Cuba, October 14–17, 2025, Proceedings, Heidelberg: Springer, 2026, p. 395-406Conference paper, Published paper (Refereed)
Abstract [en]

The subject of this work is to check how different types of music affect human emotions. While listening to music, a subjective survey and brain activity measurements were carried out using an EEG helmet. The aim is to demonstrate the impact of different music genres on emotions. The research involved a diverse group of participants of different gender and musical preferences. This had the effect of capturing a wide range of emotional responses to music. After the experiment, a relationship analysis of the respondents’ questionnaires with EEG signals was performed. The analysis revealed connections between emotions and observed brain activity. © The Author(s), under exclusive license to Springer Nature Switzerland AG 2026.

Place, publisher, year, edition, pages
Heidelberg: Springer, 2026
Series
Lecture Notes in Computer Science, ISSN 0302-9743, E-ISSN 1611-3349 ; 16328
Keywords
Music and Emotion, EEG-Based Emotion Recognition, Brain-Computer Interface (BCI)
National Category
Signal Processing
Identifiers
urn:nbn:se:hh:diva-57694 (URN)10.1007/978-3-032-11358-0_33 (DOI)2-s2.0-105029910265 (Scopus ID)978-3-032-11357-3 (ISBN)978-3-032-11358-0 (ISBN)
Conference
9th International Congress, IWAIPR: International Congress on Artificial Intelligence and Pattern Recognition, October 14–17, 2025, Varadero, Cuba
Funder
Swedish Research Council
Available from: 2025-10-30 Created: 2025-10-30 Last updated: 2026-04-22Bibliographically approved
He, L., Zhao, C., Wang, Y. & Tiwari, P. (2026). FDA-CAPMA: Federated domain adaptation with co-activation pattern and multimodal mamba for fMRI depression detection. Information Fusion, 132, Article ID 104213.
Open this publication in new window or tab >>FDA-CAPMA: Federated domain adaptation with co-activation pattern and multimodal mamba for fMRI depression detection
2026 (English)In: Information Fusion, ISSN 1566-2535, E-ISSN 1872-6305, Vol. 132, article id 104213Article in journal (Refereed) Published
Abstract [en]

Major depressive disorder is projected to become the leading contributor to mental illness by 2030. While resting-state functional magnetic resonance imaging (rs-fMRI) has emerged as a non-invasive solution for depression detection, two significant challenges remain. First, due to medical data privacy regulations and the high costs associated with acquiring the necessary equipment, individual medical institutions struggle to obtain sufficient annotated data. Second, domain shifts, caused by discrepancies in scanner parameters and acquisition protocols across multi-center datasets, significantly hinder model generalization. To address these challenges, we propose a federated domain adaptation (FDA) method that integrates co-activation patterns and a multimodal Mamba network, termed FDA-CAPMA, for fMRI-based depression detection. Specifically, a federated learning architecture ensures both physical data isolation and patient privacy through parameter aggregation. A state-space model-based Mamba network captures cross-modal correlations between fMRI time-series features and non-imaging features. Additionally, a local maximum mean discrepancy (LMMD) module aligns source and target domain distributions in both feature and prediction spaces. Extensive experiments on the largest multi-center depression dataset (Rest-meta-MDD, 1813 participants) and ABIDE dataset, our method achieves an accuracy of 67.16%, and 65.72%, respectively. This work establishes a new paradigm for privacy-preserving depression recognition. Code will be available at: https://github.com/helang818/FDA-CAPMA/ © 2026 Elsevier B.V.

Place, publisher, year, edition, pages
Amsterdam: Elsevier, 2026
Keywords
Depression, Co-activation pattern, Federated domain adaptation, Mamba
National Category
Computer Sciences
Identifiers
urn:nbn:se:hh:diva-58537 (URN)10.1016/j.inffus.2026.104213 (DOI)001697335500001 ()2-s2.0-105030338679 (Scopus ID)
Note

Funding information: This work is supported by National Natural Science Foundation of China (grant 62376215, 62236006, 62276210, 62306172, 82330043, 82471543, 62402386, 62206219, 82474666, 82105042), the Shanghai Key Laboratory of Tuina Techniques on Musculoskeletal Disorders (24dz2260200), the Three Year Action Plan for Shanghai to Further Accelerate the Inheritance, Innovation and Development of Traditional Chinese Medicine (ZY(2025–2027)-3-1-1), the Open Fund of National Engineering Laboratory for Big Data...

Available from: 2026-04-02 Created: 2026-04-02 Last updated: 2026-04-23Bibliographically approved
Liu, J., Zhang, L., Ding, Y., Tiwari, P., Guo, X. & Ding, W. (2026). Interpretable and Adaptive Graph Contrastive Learning With Information Sharing for Biomedical Link Prediction. IEEE Transactions on Emerging Topics in Computational Intelligence, 10(3), 2332-2347
Open this publication in new window or tab >>Interpretable and Adaptive Graph Contrastive Learning With Information Sharing for Biomedical Link Prediction
Show others...
2026 (English)In: IEEE Transactions on Emerging Topics in Computational Intelligence, ISSN 2471-285X, Vol. 10, no 3, p. 2332-2347Article in journal (Refereed) Published
Abstract [en]

The identification of unobserved links in drug-related biomedical networks is essential for various drug discovery applications, which is also beneficial for both disease diagnosis and treatment through exploring the underlying molecular mechanisms. However, existing solutions face significant challenges due to three main limitations, i.e., (1) lack of interpretability to provide comprehensive and reliable insights, (2) insufficient robustness and flexibility in cold-start scenarios, and (3) inadequate interaction and sharing of multi-view information. In light of this, we propose DrugXAS, an interpretable and adaptive cross-view contrastive learning framework with information sharing for biomedical link prediction. Specifically, DrugXAS has three distinctive characteristics for addressing these challenges. To solve the first problem, we propose an attention-aware augmentation scheme to provide understandable explanations of intrinsic mechanisms. To deal with the second challenge, we propose an adaptive graph updater and neighborhood sampler, which select proper neighbors according to the feedbacks from the model to improve aggregation ability. To tackle the third issue, an information sharing module with diffusion loss is proposed to incorporate chemical structures into heterogeneous relational semantics and facilitate the contrast process. Empirically, extensive experiments on seven benchmark datasets involving multi-type tasks demonstrate that DrugXAS outperforms the state-of-the-art methods in terms of precision, robustness, and interpretability. © 2025 IEEE. 

Place, publisher, year, edition, pages
Piscataway, NJ: IEEE, 2026
Keywords
Biomedical link prediction, cold-start, contrastive learning, information sharing, interpretability
National Category
Computer Sciences
Identifiers
urn:nbn:se:hh:diva-57788 (URN)10.1109/TETCI.2025.3616045 (DOI)001606748500001 ()2-s2.0-105020422631 (Scopus ID)
Note

This work was supported in part by the National Natural Science Foundation of China under Grant 62172076 and Grant 62576178, in part by the Zhejiang Provincial Natural Science Foundation of China under Grant LY23F020003, in part by the Municipal Government of Quzhou under Grant 2024D002, in part by the National Key R&D Plan of China under Grant 2024YFE0202700, and in part by the Natural Science Foundation of Jiangsu Province under Grant BK20231337.

Available from: 2025-12-11 Created: 2025-12-11 Last updated: 2026-07-08Bibliographically approved
He, L., Chen, K., Zhao, J., Wang, Y., Pei, E., Chen, H., . . . Tiwari, P. (2026). LMVD: A large-scale multimodal vlog dataset for depression detection in the wild. Information Fusion, 126(Part: B), Article ID 103632.
Open this publication in new window or tab >>LMVD: A large-scale multimodal vlog dataset for depression detection in the wild
Show others...
2026 (English)In: Information Fusion, ISSN 1566-2535, E-ISSN 1872-6305, Vol. 126, no Part: B, article id 103632Article in journal (Refereed) Published
Abstract [en]

Depression profoundly impacts multiple dimensions of an individual's life, including personal and social functioning, academic achievement, occupational productivity, and overall quality of life. With recent advancements in affective computing, deep learning technologies have been increasingly adopted to identify patterns indicative of depression. However, due to concerns over participant privacy, data in this domain remain scarce, posing significant challenges for the development of robust discriminative models for depression detection. To address this limitation, we build a Large-scale Multimodal Vlog Dataset (LMVD) for depression recognition in real-world settings. The LMVD dataset comprises 1,823 video samples, totaling approximately 214 h of content, collected from 1,475 participants across four major multimedia platforms: Sina Weibo, Bilibili, TikTok, and YouTube. In addition, we introduce a novel architecture, MDDformer, specifically designed to capture non-verbal behavioral cues associated with depressive states. Extensive experimental evaluations conducted on LMVD demonstrate the superior performance of MDDformer in depression detection tasks. We anticipate that LMVD will become a valuable benchmark resource for the research community, facilitating progress in multimodal, real-world depression recognition. The dataset and source code will be made publicly available at: https://github.com/helang818/LMVD. © 2025 Elsevier B.V., All rights reserved.

Place, publisher, year, edition, pages
Amsterdam: Elsevier, 2026
Keywords
Deep Learning, Depression Detection, Multimodal, Transformer, Vlog, Behavioral Research, Data Privacy, Human Computer Interaction, Interactive Computer Systems, Large Datasets, Learning Systems, Multimedia Systems, Academic Achievements, Deep Learning, Depression Detection, Large-scales, Multi-modal, Multiple Dimensions, Overall Quality, Quality Of Life, Transformer, Vlog
National Category
Computer Sciences
Identifiers
urn:nbn:se:hh:diva-57350 (URN)10.1016/j.inffus.2025.103632 (DOI)001619106900001 ()2-s2.0-105014021546 (Scopus ID)
Available from: 2025-09-18 Created: 2025-09-18 Last updated: 2026-02-06Bibliographically approved
Organisations
Identifiers
ORCID iD: ORCID iD iconorcid.org/0000-0002-2851-4260

Search in DiVA

Show all publications