Peer-reviewed preprints, longitudinal news datasets, and computational NLP benchmarking tools created for low-resource Nepali language processing and media intelligence.
SORTED BY DATE: NEWEST PUBLISHED / SUBMITTED FIRST
Showing 2 Peer-Reviewed & Preprint DOI Publications
109,704 unique Nepali news articles (14.3 Million tokens)Zipf's & Heap's Law linguistic validationSemantic event extraction (2015 Earthquake, COVID-19 pandemic)Linear SVM Baseline accuracy: 74.50%
Abstract & Methodology
This paper introduces Ekantipur-15Y, a long-scale longitudinal corpus of Nepali news articles spanning from 2010 to 2025. As Nepali is considered a low-resource language, the lack of a clean and temporally diverse dataset has been a barrier for the development of robust Natural Language Processing (NLP) models. We collected and cleaned 109,704 unique articles with approximately 14.3 million tokens from Ekantipur. The corpus is validated using Zipf's law confirming linguistic integrity and Heap's law demonstrating continuous growth of vocabulary without plateauing. Furthermore, the semantic analysis successfully detects major historical events in the context of Nepal (such as the 2015 earthquake and COVID-19), establishing a baseline text classification accuracy of 74.50% using Linear SVM.