作者
Duong Vu, Szániszló Szöke, Christian Wiwie, Jan Baumbach, Gianluigi Cardinali, Richard Röttger, Vincent Robert
发表日期
2014/10/30
期刊
Scientific Reports
卷号
4
期号
1
页码范围
6837
出版商
Nature Publishing Group UK
简介
With the availability of newer and cheaper sequencing methods, genomic data are being generated at an increasingly fast pace. In spite of the high degree of complexity of currently available search routines, the massive number of sequences available virtually prohibits quick and correct identification of large groups of sequences sharing common traits. Hence, there is a need for clustering tools for automatic knowledge extraction enabling the curation of large-scale databases. Current sophisticated approaches on sequence clustering are based on pairwise similarity matrices. This is impractical for databases of hundreds of thousands of sequences as such a similarity matrix alone would exceed the available memory. In this paper, a new approach called MultiLevel Clustering (MLC) is proposed which avoids a majority of sequence comparisons and therefore, significantly reduces the total runtime for clustering. An …
引用总数
2016201720182019202020212022202312213132
学术搜索中的文章