Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:Joint Chinese Word Segmentation and Part-of-speech Tagging via Two-stage Span Labeling

Dec 17, 2021

Duc-Vu Nguyen, Linh-Bao Vo, Ngoc-Linh Tran, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen

Figure 1 for Joint Chinese Word Segmentation and Part-of-speech Tagging via Two-stage Span Labeling

Figure 2 for Joint Chinese Word Segmentation and Part-of-speech Tagging via Two-stage Span Labeling

Figure 3 for Joint Chinese Word Segmentation and Part-of-speech Tagging via Two-stage Span Labeling

Figure 4 for Joint Chinese Word Segmentation and Part-of-speech Tagging via Two-stage Span Labeling

Share this with someone who'll enjoy it:

Abstract:Chinese word segmentation and part-of-speech tagging are necessary tasks in terms of computational linguistics and application of natural language processing. Many re-searchers still debate the demand for Chinese word segmentation and part-of-speech tagging in the deep learning era. Nevertheless, resolving ambiguities and detecting unknown words are challenging problems in this field. Previous studies on joint Chinese word segmentation and part-of-speech tagging mainly follow the character-based tagging model focusing on modeling n-gram features. Unlike previous works, we propose a neural model named SpanSegTag for joint Chinese word segmentation and part-of-speech tagging following the span labeling in which the probability of each n-gram being the word and the part-of-speech tag is the main problem. We use the biaffine operation over the left and right boundary representations of consecutive characters to model the n-grams. Our experiments show that our BERT-based model SpanSegTag achieved competitive performances on the CTB5, CTB6, and UD, or significant improvements on CTB7 and CTB9 benchmark datasets compared with the current state-of-the-art method using BERT or ZEN encoders.

* In Proceedings of the 35th Pacific Asia Conference on Language, Information and Computation (PACLIC 2021)

View paper on

Share this with someone who'll enjoy it:

Title:Joint Chinese Word Segmentation and Part-of-speech Tagging via Two-stage Span Labeling

Paper and Code