Back to Search Start Over

KSW: Khmer Stop Word based Dictionary for Keyword Extraction

Authors :
Thuon, Nimol
Zhang, Wangrui
Thuon, Sada
Publication Year :
2024

Abstract

This paper introduces KSW, a Khmer-specific approach to keyword extraction that leverages a specialized stop word dictionary. Due to the limited availability of natural language processing resources for the Khmer language, effective keyword extraction has been a significant challenge. KSW addresses this by developing a tailored stop word dictionary and implementing a preprocessing methodology to remove stop words, thereby enhancing the extraction of meaningful keywords. Our experiments demonstrate that KSW achieves substantial improvements in accuracy and relevance compared to previous methods, highlighting its potential to advance Khmer text processing and information retrieval. The KSW resources, including the stop word dictionary, are available at the following GitHub repository: (https://github.com/back-kh/KSWv2-Khmer-Stop-Word-based-Dictionary-for-Keyword-Extraction.git).

Details

Database :
arXiv
Publication Type :
Report
Accession number :
edsarx.2405.17390
Document Type :
Working Paper