hh.sePublications
Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Is Your LLM Outdated? A Deep Look at Temporal Generalization
The Chinese University of Hong Kong, Shenzhen, Shenzhen, China; Shenzhen Research Institute of Big Data, Shenzhen, China.
The Chinese University of Hong Kong, Shenzhen, Shenzhen, China.
The Chinese University of Hong Kong, Shenzhen, Shenzhen, China.
The Chinese University of Hong Kong, Shenzhen, Shenzhen, China.
Show others and affiliations
2025 (English)In: Proceedings of the 2025 Annual Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies - Long Papers, NAACL-HLT 2025 / [ed] Chiruzzo L.; Ritter A.; Wang L, Albuquerque: Association for Computational Linguistics, 2025, Vol. 1, p. 7433-7457Conference paper, Published paper (Refereed)
Abstract [en]

The rapid advancement of Large Language Models (LLMs) has led to the development of benchmarks that consider temporal dynamics, however, there remains a gap in understanding how well these models can generalize across temporal contexts due to the inherent dynamic nature of language and information. This paper introduces the concept of temporal generalization in LLMs, including bias in past and future generalizations. Then we introduce FreshBench, a new evaluation framework that employs fresh text and event prediction for assessing LLMs' temporal adaptability, ensuring the evaluation process free from data leakage and subjective bias. The experiment shows significant temporal biases and a decline in performance over time. Our findings reveal that powerful models, while initially superior, tend to decline more rapidly in future generalization. Additionally, powerful open-source models demonstrate better long-term adaptability compared to their closed-source counterparts. Our code is available at https://github.com/FreedomIntelligence/FreshBench. © 2025 Association for Computational Linguistics.

Place, publisher, year, edition, pages
Albuquerque: Association for Computational Linguistics, 2025. Vol. 1, p. 7433-7457
National Category
Computer and Information Sciences
Identifiers
URN: urn:nbn:se:hh:diva-58221DOI: 10.18653/v1/2025.naacl-long.381Scopus ID: 2-s2.0-105027416237ISBN: 9798891761896 (electronic)OAI: oai:DiVA.org:hh-58221DiVA, id: diva2:2034102
Conference
2025 Annual Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2025, Albuquerque, New Mexico, 29 April - 4 May, 2025
Available from: 2026-01-30 Created: 2026-01-30 Last updated: 2026-01-30Bibliographically approved

Open Access in DiVA

No full text in DiVA

Other links

Publisher's full textScopus

Authority records

Tiwari, Prayag

Search in DiVA

By author/editor
Tiwari, PrayagWang, Benyou
By organisation
School of Information Technology
Computer and Information Sciences

Search outside of DiVA

GoogleGoogle Scholar

doi
isbn
urn-nbn

Altmetric score

doi
isbn
urn-nbn
Total: 31 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf