N-gram and local context analysis for Persian text retrieval

A Aleahmad, P Hakimian, F Mahdikhani… - … symposium on signal …, 2007 - ieeexplore.ieee.org
A Aleahmad, P Hakimian, F Mahdikhani, F Oroumchian
2007 9th international symposium on signal processing and its …, 2007ieeexplore.ieee.org
The Persian language is one of the languages in Middle-East, so there are significant
amount of Persian documents available on the Web. But there are relatively few studies on
retrieval of Persian documents in the literature. In this experimental study, we assessed term
and N-gram based vector space model and a query expansion method, namely, local
context analysis using different weighting schemes on a realistic corpus containing 160000+
news articles. Then we compared our results with previous works reported on Persian …
The Persian language is one of the languages in Middle-East, so there are significant amount of Persian documents available on the Web. But there are relatively few studies on retrieval of Persian documents in the literature. In this experimental study, we assessed term and N-gram based vector space model and a query expansion method, namely, local context analysis using different weighting schemes on a realistic corpus containing 160000+ news articles. Then we compared our results with previous works reported on Persian language. Our experimental results show that among the assessed methods, 4-gram based vector space model with Lnu.ltu weighting scheme has acceptable performance and Local context analysis has the best performance for Persian text retrieval so far.
ieeexplore.ieee.org
以上显示的是最相近的搜索结果。 查看全部搜索结果