Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Abdelmonaime Lachkar

A new keyphrases extraction method based on suffix tree data structure for arabic documents clustering

Jan 22, 2014

Issam Sahmoudi, Hanane Froud, Abdelmonaime Lachkar

Figure 1 for A new keyphrases extraction method based on suffix tree data structure for arabic documents clustering

Figure 2 for A new keyphrases extraction method based on suffix tree data structure for arabic documents clustering

Abstract:Document Clustering is a branch of a larger area of scientific study known as data mining .which is an unsupervised classification using to find a structure in a collection of unlabeled data. The useful information in the documents can be accompanied by a large amount of noise words when using Full Text Representation, and therefore will affect negatively the result of the clustering process. So it is with great need to eliminate the noise words and keeping just the useful information in order to enhance the quality of the clustering results. This problem occurs with different degree for any language such as English, European, Hindi, Chinese, and Arabic Language. To overcome this problem, in this paper, we propose a new and efficient Keyphrases extraction method based on the Suffix Tree data structure (KpST), the extracted Keyphrases are then used in the clustering process instead of Full Text Representation. The proposed method for Keyphrases extraction is language independent and therefore it may be applied to any language. In this investigation, we are interested to deal with the Arabic language which is one of the most complex languages. To evaluate our method, we conduct an experimental study on Arabic Documents using the most popular Clustering approach of Hierarchical algorithms: Agglomerative Hierarchical algorithm with seven linkage techniques and a variety of distance functions and similarity measures to perform Arabic Document Clustering task. The obtained results show that our method for extracting Keyphrases increases the quality of the clustering results. We propose also to study the effect of using the stemming for the testing dataset to cluster it with the same documents clustering techniques and similarity/distance measures.

* International Journal of Database Management Systems ( IJDMS ) Vol.5, No.6, December 2013
* 17 pages, 3 figures

Via

Access Paper or Ask Questions

Arabic text summarization based on latent semantic analysis to enhance arabic documents clustering

Feb 06, 2013

Hanane Froud, Abdelmonaime Lachkar, Said Alaoui Ouatik

Figure 1 for Arabic text summarization based on latent semantic analysis to enhance arabic documents clustering

Figure 2 for Arabic text summarization based on latent semantic analysis to enhance arabic documents clustering

Figure 3 for Arabic text summarization based on latent semantic analysis to enhance arabic documents clustering

Figure 4 for Arabic text summarization based on latent semantic analysis to enhance arabic documents clustering

Abstract:Arabic Documents Clustering is an important task for obtaining good results with the traditional Information Retrieval (IR) systems especially with the rapid growth of the number of online documents present in Arabic language. Documents clustering aim to automatically group similar documents in one cluster using different similarity/distance measures. This task is often affected by the documents length, useful information on the documents is often accompanied by a large amount of noise, and therefore it is necessary to eliminate this noise while keeping useful information to boost the performance of Documents clustering. In this paper, we propose to evaluate the impact of text summarization using the Latent Semantic Analysis Model on Arabic Documents Clustering in order to solve problems cited above, using five similarity/distance measures: Euclidean Distance, Cosine Similarity, Jaccard Coefficient, Pearson Correlation Coefficient and Averaged Kullback-Leibler Divergence, for two times: without and with stemming. Our experimental results indicate that our proposed approach effectively solves the problems of noisy information and documents length, and thus significantly improve the clustering performance.

* International Journal of Data Mining & Knowledge Management Process (IJDKP)- 2013

Via

Access Paper or Ask Questions