Knowledgeable Salient Span Mask for Enhancing Language Models as Knowledge Base
Wang, Cunxiang, Luo, Fuli, Li, Yanyang, Xu, Runxin, Huang, Fei, Zhang, Yue
–arXiv.org Artificial Intelligence
Pre-trained language models (PLMs) like BERT have made significant progress in various downstream NLP tasks. However, by asking models to do cloze-style tests, recent work finds that PLMs are short in acquiring knowledge from unstructured text. To understand the internal behaviour of PLMs in retrieving knowledge, we first define knowledge-baring (K-B) tokens and knowledge-free (K-F) tokens for unstructured text and ask professional annotators to label some samples manually. Then, we find that PLMs are more likely to give wrong predictions on K-B tokens and attend less attention to those tokens inside the self-attention module. Based on these observations, we develop two solutions to help the model learn more knowledge from unstructured text in a fully self-supervised manner. Experiments on knowledge-intensive tasks show the effectiveness of the proposed methods. To our best knowledge, we are the first to explore fully self-supervised learning of knowledge in continual pre-training.
arXiv.org Artificial Intelligence
Oct-11-2023
- Country:
- North America
- Dominican Republic (0.04)
- United States
- Maryland (0.04)
- Minnesota > Hennepin County
- Minneapolis (0.14)
- Europe > Middle East
- Republic of Türkiye > Istanbul Province > Istanbul (0.04)
- Asia
- China > Hong Kong (0.04)
- Middle East > Republic of Türkiye
- Istanbul Province > Istanbul (0.04)
- Japan > Kyūshū & Okinawa
- Kyūshū > Miyazaki Prefecture > Miyazaki (0.04)
- Africa
- Middle East > Morocco (0.04)
- Kenya (0.04)
- North America
- Genre:
- Research Report (0.64)
- Technology: