Do Transformers Need Deep Long-Range Memory

Rae, Jack W., Razavi, Ali

arXiv.org Machine Learning 

Do Transformers Need Deep Long-Range Memory? Daniluk et al. (2017) observed that an LSTM We do this with a selective Transformer-XL (Dai et al., 2019), a Transformer memorisation process; most of the finer details of variant specialised for long-range sequence modelling the text are quickly forgotten and we retain a relatively via the introduction of a cache of past activations, compact representation of the book's details. obtained state-of-the-art results in the four Early models of natural language used recurrent major LM benchmarks -- PTB (Mikolov et al., neural networks (RNNs) such as the Long Short-2010), LM1B (Chelba et al., 2013), Enwik8 (Hutter, Term Memory (Hochreiter and Schmidhuber, 1997) 2012), and WikiText (Merity et al., 2016).

Duplicate Docs Excel Report

Title
None found

Similar Docs  Excel Report  more

TitleSimilaritySource
None found