Turning transformer attention weights into zero-shot sequence labelers
Bujel, Kamil, Yannakoudakis, Helen, Rei, Marek
–arXiv.org Artificial Intelligence
We demonstrate how transformer-based models can be redesigned in order to capture inductive biases across tasks on different granularities and perform inference in a zero-shot manner. Specifically, we show how sentence-level transformers can be modified into effective sequence labelers at the token level without any direct supervision. We compare against a range of diverse and previously proposed methods for generating token-level labels, and present a simple yet effective modified attention layer that significantly advances the current state of the art.
arXiv.org Artificial Intelligence
Mar-26-2021
- Country:
- North America > United States
- Oregon > Multnomah County
- Portland (0.04)
- New Mexico > Santa Fe County
- Santa Fe (0.04)
- Minnesota > Hennepin County
- Minneapolis (0.15)
- Louisiana > Orleans Parish
- New Orleans (0.04)
- California > San Diego County
- San Diego (0.04)
- Oregon > Multnomah County
- Europe
- Germany > Berlin (0.04)
- United Kingdom > England
- Greater London > London (0.04)
- Sweden > Uppsala County
- Uppsala (0.04)
- Italy > Tuscany
- Florence (0.05)
- Belgium > Brussels-Capital Region
- Brussels (0.04)
- Asia > China
- Hong Kong (0.04)
- North America > United States
- Genre:
- Research Report (1.00)
- Technology: