Are BERT Features InterBERTible?

#artificialintelligence 

We've come a long way in the word embedding space since the introduction of Word2Vec (Mikolov et. These days, it seems that every single machine learning practitioner can recite the "king minus man plus woman equals queen" mantra. In present, these interpretable word embeddings have become an essential part in many deep-learning based NLP systems. Earlier last October, Google AI introduced BERT: Bidirectional Encoder Representations from Transformers (paper, source). Seemingly, the researchers at Google have done it again: they've come up with a model to learn contextual word representations that redefined the state of the art for 11 NLP tasks, 'even surpassing human performance in the challenging area of question answering'.