查看文章

ieee.org 中的 [PDF]

Compilation, analysis and application of a comprehensive Bangla Corpus KUMono

作者

Aysha Akther, Md Shymon Islam, Hafsa Sultana, AKZ Rasel Rahman, Sujana Saha, Kazi Masudul Alam, Rameswar Debnath

发表日期

2022/8/1

期刊

IEEE Access

卷号

页码范围

79999-80014

出版商

IEEE

简介

Research in Natural Language Processing (NLP) and computational linguistics highly depends on a good quality representative corpus of any specific language. Bangla is one of the most spoken languages in the world but Bangla NLP research is in its early stage of development due to the lack of quality public corpus. This article describes the detailed compilation methodology of a comprehensive monolingual Bangla corpus, KUMono ( K hulna U niversity Mono lingual corpus). The newly developed corpus consists of more than 350 million word tokens and more than one million unique tokens from 18 major text categories of online Bangla websites. We have conducted several word-level and character-level linguistic phenomenon analyses based on empirical studies of the developed corpus. The corpus follows Zipf’s curve and hapax legomena rule. The quality of the corpus is also assessed by analyzing and …

引用总数

被引用次数：9

202320246 3

学术搜索中的文章

Compilation, analysis and application of a comprehensive Bangla Corpus KUMono

A Akther, MS Islam, H Sultana, AKZR Rahman, S Saha… - IEEE Access, 2022

被引用次数：9 相关文章所有 5 个版本