Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Kaixu Zhang

Reduce Meaningless Words for Joint Chinese Word Segmentation and Part-of-speech Tagging

May 25, 2013

Kaixu Zhang, Maosong Sun

Figure 1 for Reduce Meaningless Words for Joint Chinese Word Segmentation and Part-of-speech Tagging

Figure 2 for Reduce Meaningless Words for Joint Chinese Word Segmentation and Part-of-speech Tagging

Figure 3 for Reduce Meaningless Words for Joint Chinese Word Segmentation and Part-of-speech Tagging

Figure 4 for Reduce Meaningless Words for Joint Chinese Word Segmentation and Part-of-speech Tagging

Abstract:Conventional statistics-based methods for joint Chinese word segmentation and part-of-speech tagging (S&T) have generalization ability to recognize new words that do not appear in the training data. An undesirable side effect is that a number of meaningless words will be incorrectly created. We propose an effective and efficient framework for S&T that introduces features to significantly reduce meaningless words generation. A general lexicon, Wikepedia and a large-scale raw corpus of 200 billion characters are used to generate word-based features for the wordhood. The word-lattice based framework consists of a character-based model and a word-based model in order to employ our word-based features. Experiments on Penn Chinese treebank 5 show that this method has a 62.9% reduction of meaningless word generation in comparison with the baseline. As a result, the F1 measure for segmentation is increased to 0.984.

Via

Access Paper or Ask Questions

Binary Tree based Chinese Word Segmentation

May 17, 2013

Kaixu Zhang, Can Wang, Maosong Sun

Figure 1 for Binary Tree based Chinese Word Segmentation

Figure 2 for Binary Tree based Chinese Word Segmentation

Figure 3 for Binary Tree based Chinese Word Segmentation

Figure 4 for Binary Tree based Chinese Word Segmentation

Abstract:Chinese word segmentation is a fundamental task for Chinese language processing. The granularity mismatch problem is the main cause of the errors. This paper showed that the binary tree representation can store outputs with different granularity. A binary tree based framework is also designed to overcome the granularity mismatch problem. There are two steps in this framework, namely tree building and tree pruning. The tree pruning step is specially designed to focus on the granularity problem. Previous work for Chinese word segmentation such as the sequence tagging can be easily employed in this framework. This framework can also provide quantitative error analysis methods. The experiments showed that after using a more sophisticated tree pruning function for a state-of-the-art conditional random field based baseline, the error reduction can be up to 20%.

Via

Access Paper or Ask Questions