Neural Chinese Word Segmentation with Lexicon and Unlabeled Data via Posterior Regularization

Liu, Junxin, Wu, Fangzhao, Wu, Chuhan, Huang, Yongfeng, Xie, Xing

arXiv.org Machine Learning 

Existing methods for CWS usually rely on a large In recent years, neural network based methods have been widely number of labeled sentences to train word segmentation models, used for CWS [1, 17, 23, 26]. Most of these methods model CWS as which are expensive and time-consuming to annotate. Luckily, the a sequence labeling problem [22, 30], and utilize neural networks unlabeled data is usually easy to collect and many high-quality to learn the hidden character features [2, 32]. For example, Chen et Chinese lexicons are off-the-shelf, both of which can provide useful al. [2] used LSTM [8] to learn character features by capturing the information for CWS. In this paper, we propose a neural approach global information of sentence. Peng and Dredze [18] proposed to for Chinese word segmentation which can exploit both lexicon use LSTM for character feature learning and CRF [9] for character and unlabeled data. Our approach is based on a variant of posterior label decoding. However, these methods usually rely on a large regularization algorithm, and the unlabeled data and lexicon number of labeled sentences to train word segmentation models, are incorporated into model training as indirect supervision by which are expensive and time-consuming to annotate.

Duplicate Docs Excel Report

Title
None found

Similar Docs  Excel Report  more

TitleSimilaritySource
None found