Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:A Subword Guided Neural Word Segmentation Model for Sindhi

Dec 30, 2020

Wazir Ali, Jay Kumar, Zenglin Xu, Congjian Luo, Junyu Lu, Junming Shao, Rajesh Kumar, Yazhou Ren

Figure 1 for A Subword Guided Neural Word Segmentation Model for Sindhi

Figure 2 for A Subword Guided Neural Word Segmentation Model for Sindhi

Figure 3 for A Subword Guided Neural Word Segmentation Model for Sindhi

Figure 4 for A Subword Guided Neural Word Segmentation Model for Sindhi

Share this with someone who'll enjoy it:

Abstract:Deep neural networks employ multiple processing layers for learning text representations to alleviate the burden of manual feature engineering in Natural Language Processing (NLP). Such text representations are widely used to extract features from unlabeled data. The word segmentation is a fundamental and inevitable prerequisite for many languages. Sindhi is an under-resourced language, whose segmentation is challenging as it exhibits space omission, space insertion issues, and lacks the labeled corpus for segmentation. In this paper, we investigate supervised Sindhi Word Segmentation (SWS) using unlabeled data with a Subword Guided Neural Word Segmenter (SGNWS) for Sindhi. In order to learn text representations, we incorporate subword representations to recurrent neural architecture to capture word information at morphemic-level, which takes advantage of Bidirectional Long-Short Term Memory (BiLSTM), self-attention mechanism, and Conditional Random Field (CRF). Our proposed SGNWS model achieves an F1 value of 98.51% without relying on feature engineering. The empirical results demonstrate the benefits of the proposed model over the existing Sindhi word segmenters.

* Journal Paper, 16 pages

View paper on

Share this with someone who'll enjoy it:

Title:A Subword Guided Neural Word Segmentation Model for Sindhi

Paper and Code