BitFit: Simple Parameter-efficient Fine-tuning for Transformer-based Masked Language-models
Zaken, Elad Ben, Ravfogel, Shauli, Goldberg, Yoav
–arXiv.org Artificial Intelligence
We introduce BitFit, a sparse-finetuning method where only the bias-terms of the model (or a subset of them) are being modified. We show that with small-to-medium training data, applying BitFit on pre-trained BERT models is competitive with (and sometimes better than) fine-tuning the entire model. For larger data, the method is competitive with other sparse fine-tuning methods. Besides their practical utility, these findings are relevant for the question of understanding the commonly-used process of finetuning: they support the hypothesis that finetuning is mainly about exposing knowledge induced by language-modeling training, rather than learning new task-specific linguistic knowledge.
arXiv.org Artificial Intelligence
Sep-5-2022
- Country:
- Asia > Singapore (0.04)
- Oceania > Australia
- North America > United States
- Louisiana > Orleans Parish
- New Orleans (0.04)
- California > Los Angeles County
- Long Beach (0.04)
- Louisiana > Orleans Parish
- Europe
- France (0.04)
- United Kingdom > England
- Cambridgeshire > Cambridge (0.04)
- Romania > Sud - Muntenia Development Region
- Giurgiu County > Giurgiu (0.04)
- Genre:
- Research Report (0.82)
- Technology: