作者
R Priyadarshini, Latha Tamilselvan
发表日期
2018
期刊
International Journal of Networking and Virtual Organisations
卷号
19
期号
2-4
页码范围
321-340
出版商
Inderscience Publishers (IEL)
简介
Today's era is rather called big data era, data starts growing from different sources of web and such scalable data is very hard to manage with the existing frameworks and technologies. Wikipedia is a content management system where the article posted has a number of source documents. Perhaps, it is very difficult to search an exact relevant document for selected content in Wikipedia article as it has too many sources such as primary, secondary and tertiary. In order to search and retrieve relevant document in the growing content and references, clustering of documents using similarity analysis is very much essential. The existing system offers a clustering technique based on term and inverse term frequency (TfDf) scoring method. This work proposes a new clustering method for distributed framework called semantic agglomerative hierarchical (SHA) clustering algorithm. The performance testing, evaluation is …
引用总数
学术搜索中的文章