BERT and PALs: Projected Attention Layers for Efficient Adaptation in Multi-Task Learning
Stickland, Asa Cooper, Murray, Iain
Multi-task learning allows the sharing of useful information between multiple related tasks. In natural language processing several recent approaches havesuccessfully leveraged unsupervised pre-training on large amounts of data to perform well on various tasks, such as those in the GLUE benchmark (Wang et al., 2018a). These results are based on fine-tuning on each task separately. Weexplore the multi-task learning setting for the recent BERT (Devlin et al., 2018) model on the GLUE benchmark, and how to best add task-specific parameters to a pre-trained BERT network, with a high degree of parameter sharing between tasks. We introduce new adaptation modules, PALsor'projected attention layers', which use a low-dimensional multi-head attention mechanism, basedon the idea that it is important to include layers with inductive biases useful for the input domain. By using PALs in parallel with BERT layers, we match the performance of finetuned BERTon the GLUE benchmark with 7 times fewer parameters, and obtain state-of-theart resultson the Recognizing Textual Entailment dataset.
Feb-7-2019
- Country:
- Europe > Romania > Sud - Muntenia Development Region > Giurgiu County > Giurgiu (0.04)
- Genre:
- Research Report (1.00)
- Technology: