Knowledgeable or Educated Guess? Revisiting Language Models as Knowledge Bases

Cao, Boxi, Lin, Hongyu, Han, Xianpei, Sun, Le, Yan, Lingyong, Liao, Meng, Xue, Tong, Xu, Jin

Jun-16-2021–arXiv.org Artificial Intelligence

Previous literatures show that pre-trained masked language models (MLMs) such as BERT can achieve competitive factual knowledge extraction performance on some datasets, indicating that MLMs can potentially be a reliable knowledge source. In this paper, we conduct a rigorous study to explore the underlying predicting mechanisms of MLMs over different extraction paradigms. By investigating the behaviors of MLMs, we find that previous decent performance mainly owes to the biased prompts which overfit dataset artifacts. Furthermore, incorporating illustrative cases and external contexts improve knowledge prediction mainly due to entity type guidance and golden answer leakage. Our findings shed light on the underlying predicting mechanisms of MLMs, and strongly question the previous conclusion that current MLMs can potentially serve as reliable factual knowledge bases.

computational linguistic, prediction, prediction distribution, (14 more...)

arXiv.org Artificial Intelligence

Jun-16-2021

arXiv.org PDF

Add feedback

Country:
- North America
  - United States
    - Hawaii (0.04)
    - New York > New York County
      - New York City (0.04)
    - Minnesota > Hennepin County
      - Minneapolis (0.28)
    - Louisiana > Orleans Parish
      - New Orleans (0.04)
    - Illinois > Cook County
      - Chicago (0.05)
    - California > San Diego County
      - San Diego (0.04)
  - Canada
    - Quebec > Montreal (0.04)
    - British Columbia > Metro Vancouver Regional District
      - Vancouver (0.04)
- Europe
  - Russia > Central Federal District
    - Moscow Oblast > Moscow (0.04)
  - Italy > Tuscany
    - Florence (0.04)
- Asia
  - Middle East > Jordan (0.04)
  - Japan
    - Kyūshū & Okinawa > Kyūshū
      - Miyazaki Prefecture > Miyazaki (0.04)
    - Honshū > Kantō
      - Tokyo Metropolis Prefecture > Tokyo (0.04)
  - China
    - Hong Kong (0.04)
    - Beijing > Beijing (0.04)

Genre:
- Research Report > New Finding (0.88)

Technology:
- Information Technology > Artificial Intelligence
  - Natural Language (1.00)
  - Representation & Reasoning > Expert Systems (0.70)

Duplicate Docs Excel Report

Title
None found

Similar Docs Excel Report more

Title	Similarity	Source
None found