hh.sePublications
Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Neural Network Optimization Reimagined: Decoupled Techniques for Scratch and Fine-Tuning
Institute of Semiconductors, Beijing, China.ORCID iD: 0000-0001-7897-1673
Nanyang Technological University, Singapore City, Singapore.
Montreal Institute for Learning Algorithms, Montreal, Canada.
University of Science and Technology of China, Hefei, China.
Show others and affiliations
2026 (English)In: IEEE Transactions on Pattern Analysis and Machine Intelligence, ISSN 0162-8828, E-ISSN 1939-3539Article in journal (Refereed) Epub ahead of print
Abstract [en]

With the accumulation of resources in the era of big data and the rise of pre-trained models in deep learning, optimizing neural networks for various tasks often involves different strategies for fine-tuning pre-trained models versus training from scratch. However, existing optimizers primarily focus on reducing the loss function by updating model parameters, without fully addressing the unique demands of these two major paradigms. In this paper, we propose DualOpt, a novel approach that decouples optimization techniques specifically tailored for these distinct training scenarios. For training from scratch, we introduce real-time layer-wise weight decay, designed to enhance both convergence and generalization by aligning with the characteristics of weight updates and network architecture. For more importantly fine-tuning, we integrate weight rollback with the optimizer, incorporating a rollback term into each weight update step. This ensures consistency in the weight distribution between upstream and downstream models, effectively mitigating knowledge forgetting and improving fine-tuning performance. Additionally, we extend the layer-wise weight decay to dynamically adjust the rollback levels across layers, adapting to the varying demands of different downstream tasks. Extensive experiments across diverse tasks, including image classification, object detection, semantic segmentation, and instance segmentation, demonstrate the broad applicability and state-of-the-art performance of DualOpt. © 2026 IEEE.

Place, publisher, year, edition, pages
Piscataway: IEEE, 2026.
Keywords [en]
fine-tuning, layer-wise penalty, Neural network optimization, training from scratch, weight rollback
National Category
Computer graphics and computer vision Artificial Intelligence Security, Privacy and Cryptography
Identifiers
URN: urn:nbn:se:hh:diva-58936DOI: 10.1109/TPAMI.2026.3683792PubMedID: 41979964Scopus ID: 2-s2.0-105036738768OAI: oai:DiVA.org:hh-58936DiVA, id: diva2:2059985
Note

Funding information: This work is funded by Beijing Natural Science Foundation (No. L233036) and National Natural Science Foundation of China (No. 62373343).

Available from: 2026-05-13 Created: 2026-05-13 Last updated: 2026-05-13Bibliographically approved

Open Access in DiVA

No full text in DiVA

Other links

Publisher's full textPubMedScopus

Authority records

Tiwari, Prayag

Search in DiVA

By author/editor
Ning, XinHe, FengTiwari, PrayagLiu, Xinwang
By organisation
School of Information Technology
In the same journal
IEEE Transactions on Pattern Analysis and Machine Intelligence
Computer graphics and computer visionArtificial IntelligenceSecurity, Privacy and Cryptography

Search outside of DiVA

GoogleGoogle Scholar

doi
pubmed
urn-nbn

Altmetric score

doi
pubmed
urn-nbn
Total: 8 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf