hh.sePublications
Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
From Courtroom to Corpora: Building a Name Entity Corpus for Urdu Legal Texts
Halmstad University, School of Information Technology.ORCID iD: 0000-0002-4826-7746
Riphah International University, Islamabad, Pakistan.
Halmstad University, School of Information Technology, Center for Applied Intelligent Systems Research (CAISR).ORCID iD: 0000-0002-7796-5201
2025 (English)In: Natural Language Processing in the Generative AI Era: PROCEEDINGS / [ed] Galia Angelova; Maria Kunilovskaya; Marie Escribe; Ruslan Mitkov, Shoumen: INCOMA Ltd , 2025, p. 1396-1405Conference paper, Published paper (Refereed)
Abstract [en]

This study explores the effectiveness of transformer-based models for Named Entity Recognition (NER) in Urdu legal documents, a critical task in low-resource language processing. Given the specialized terminology and complex syntax of the legal texts, accurate entity recognition in Urdu remains a challenge. We developed a legal Urdu dataset that contains 117,500 documents, generated synthetically from 47 different types of legal documents, and evaluated three BERT-based models. XLMRoBERTa, mBERT, and DistilBERT were analyzed by analyzing their performance on an annotated Urdu legal data set. mBERT demonstrated superior accuracy (0.999), and its F1 score (0.975) outperforms XLMRoBERTa and DistilBERT, highlighting its robustness in recognizing entities within low-resource languages. To ensure the privacy of personal identifiers, all documents are anonymized. The dataset for this study is publicly hosted on HuggingFace. 

Place, publisher, year, edition, pages
Shoumen: INCOMA Ltd , 2025. p. 1396-1405
Series
International conference Recent advances in natural language processing, ISSN 1313-8502, E-ISSN 2603-2813
Keywords [en]
Named Entity Recognition, Low Resource Languages, Synthetic Data, DistilBERT, XLM-RoBERTa, mBERT
National Category
Computer and Information Sciences
Identifiers
URN: urn:nbn:se:hh:diva-58445DOI: 10.26615/978-954-452-098-4-161OAI: oai:DiVA.org:hh-58445DiVA, id: diva2:2038885
Conference
The International Conference on Recent Advances in Natural Language Processing (RANLP), Varna, Bulgaria 8–10 September, 2025
Available from: 2026-02-16 Created: 2026-02-16 Last updated: 2026-02-17Bibliographically approved

Open Access in DiVA

No full text in DiVA

Other links

Publisher's full text

Authority records

Adeel, ZafarNowaczyk, Sławomir

Search in DiVA

By author/editor
Adeel, ZafarNowaczyk, Sławomir
By organisation
School of Information TechnologyCenter for Applied Intelligent Systems Research (CAISR)
Computer and Information Sciences

Search outside of DiVA

GoogleGoogle Scholar

doi
urn-nbn

Altmetric score

doi
urn-nbn
Total: 70 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf