Do Transformers Need Deep Long-Range Memory
Do Transformers Need Deep Long-Range Memory? Daniluk et al. (2017) observed that an LSTM We do this with a selective Transformer-XL (Dai et al., 2019), a Transformer memorisation process; most of the finer details of variant specialised for long-range sequence modelling the text are quickly forgotten and we retain a relatively via the introduction of a cache of past activations, compact representation of the book's details. obtained state-of-the-art results in the four Early models of natural language used recurrent major LM benchmarks -- PTB (Mikolov et al., neural networks (RNNs) such as the Long Short-2010), LM1B (Chelba et al., 2013), Enwik8 (Hutter, Term Memory (Hochreiter and Schmidhuber, 1997) 2012), and WikiText (Merity et al., 2016).
Jul-7-2020
- Country:
- Europe
- United Kingdom > England
- Greater London > London (0.05)
- Italy > Calabria
- Catanzaro Province > Catanzaro (0.04)
- United Kingdom > England
- Europe
- Genre:
- Research Report (0.64)
- Technology: