Adaptive Boosting LLMs for Text ClassificationShow others and affiliations
2026 (English)In: IEEE Transactions on Neural Networks and Learning Systems, ISSN 2162-237X, E-ISSN 2162-2388, Vol. 37, no 6, p. 2770-2781Article in journal (Refereed) Published
Abstract [en]
With large-scale language models demonstrating superior capabilities in a wide range of downstream natural language processing tasks, the future trajectory of research in the field of text categorization faces increasing uncertainty. In this evolving paradigm of open-ended language modeling, where task delimitations are increasingly blurred, a pressing question arises: to what extent has text classification advanced under the full potential of large language model (LLM)? To address this pivotal inquiry, we introduce recurrent generative pre-trained transformer (RGPT), an adaptive boosting framework meticulously designed to craft a dedicated LLM for text classification. RGPT constructs a sequence of base learners by dynamically modulating the training data distribution and iteratively fine-tuning LLMs. These base learners are then progressively integrated, leveraging historical prediction trajectories to form a highly specialized text classification model. Extensive empirical evaluations demonstrate that RGPT surpasses eight state-of-the-art pretrained language models and seven cutting-edge LLMs across four benchmark datasets, achieving an average performance gain of 2.90%. © 2012 IEEE.
Place, publisher, year, edition, pages
Piscataway, NJ: IEEE, 2026. Vol. 37, no 6, p. 2770-2781
Keywords [en]
Boosting, large language model, natural language processing, text classification
National Category
Natural Language Processing
Identifiers
URN: urn:nbn:se:hh:diva-58202DOI: 10.1109/TNNLS.2025.3639613ISI: 001663478400001PubMedID: 41525530Scopus ID: 2-s2.0-105027441383OAI: oai:DiVA.org:hh-58202DiVA, id: diva2:2034089
Note
Funding information: This work was supported in part by the National Social Science Fund of China under Grant 25AYY001; in part by the Collaborative Research with World-Leading Research Groups Scheme in The Hong Kong Polytechnic University under Project G-SACF; in part by the General Research Fund through Hong Kong Research Grants Council under Project 15218521; in part by the Natural Science Foundation of Henan Province of China under Grant 242300421412; and in part by the Foundation of Key Laboratory of Dependable Service Computing in CyberPhysical-Society (Ministry of Education), Chongqing University under Grant CPSDSC202103.
2026-01-302026-01-302026-07-06Bibliographically approved