Multi-CLS BERT: An Efficient Alternative to Traditional Ensembling
Chang, Haw-Shiuan, Sun, Ruei-Yao, Ricci, Kathryn, McCallum, Andrew
–arXiv.org Artificial Intelligence
Ensembling BERT models often significantly improves accuracy, but at the cost of significantly more computation and memory footprint. In this work, we propose Multi-CLS BERT, a novel ensembling method for CLS-based prediction tasks that is almost as efficient as a single BERT model. Multi-CLS BERT uses multiple CLS tokens with a parameterization and objective that encourages their diversity. Thus instead of fine-tuning each BERT model in an ensemble (and running them all at test time), we need only fine-tune our single Multi-CLS BERT model (and run the one model at test time, ensembling just the multiple final CLS embeddings). To test its effectiveness, we build Multi-CLS BERT on top of a state-of-the-art pretraining method for BERT (Aroca-Ouellette and Rudzicz, 2020). In experiments on GLUE and SuperGLUE we show that our Multi-CLS BERT reliably improves both overall accuracy and confidence estimation. When only 100 training samples are available in GLUE, the Multi-CLS BERT_Base model can even outperform the corresponding BERT_Large model. We analyze the behavior of our Multi-CLS BERT, showing that it has many of the same characteristics and behavior as a typical BERT 5-way ensemble, but with nearly 4-times less computation and memory.
arXiv.org Artificial Intelligence
May-20-2023
- Country:
- South America > Chile
- North America
- Mexico (0.04)
- Dominican Republic (0.04)
- United States
- Utah (0.04)
- Virginia (0.04)
- Texas > Travis County
- Austin (0.04)
- Minnesota > Hennepin County
- Minneapolis (0.14)
- New York
- New York County > New York City (0.27)
- Richmond County > New York City (0.14)
- Queens County > New York City (0.04)
- Kings County > New York City (0.04)
- Bronx County > New York City (0.04)
- Louisiana > Orleans Parish
- New Orleans (0.04)
- Massachusetts
- Norfolk County (0.04)
- Hampshire County > Amherst (0.04)
- Pennsylvania > Erie County
- Erie (0.04)
- California
- San Diego County > San Diego (0.04)
- Monterey County > Monterey (0.04)
- Los Angeles County
- Los Angeles (0.04)
- Long Beach (0.04)
- Canada > British Columbia
- Europe
- Austria (0.04)
- Monaco (0.04)
- France (0.04)
- United Kingdom
- Romania > Sud - Muntenia Development Region
- Giurgiu County > Giurgiu (0.04)
- Italy > Tuscany
- Florence (0.04)
- Asia
- Middle East > Iraq (0.04)
- China > Hong Kong (0.04)
- Myanmar > Yangon Region
- Yangon (0.04)
- Japan > Kyūshū & Okinawa
- Kyūshū > Nagasaki Prefecture > Nagasaki (0.04)
- Africa
- Middle East > Egypt (0.04)
- Nigeria (0.04)
- Ethiopia > Addis Ababa
- Addis Ababa (0.04)
- Genre:
- Research Report > New Finding (0.45)
- Industry:
- Media > Film (1.00)
- Law Enforcement & Public Safety (1.00)
- Law (1.00)
- Health & Medicine (1.00)
- Leisure & Entertainment > Sports (0.92)
- Government > Regional Government
- Technology: