precompute
Transformer tricks: Precomputing the first layer
This micro-paper describes a trick to speed up inference of transformers with RoPE (such as LLaMA, Mistral, PaLM, and Gemma). For these models, a large portion of the first transformer layer can be precomputed, which results in slightly lower latency and lower cost-per-token. Because this trick optimizes only one layer, the relative savings depend on the total number of layers. For example, the maximum savings for a model with only 4 layers (such as Whisper tiny) is limited to 25%, while a 32-layer model is limited to 3% savings. See https://github.com/OpenMachine-ai/transformer-tricks for code and more transformer tricks.
When shuffling large arrays, how much time can be attributed to random number generation?
It is well known that contemporary computers don't like to randomly access data in an unpredictible manner in memory. However, not all forms of random accesses are equally harmful. Suppose that the array is large. Take an array made of 100 million elements. It far exceeds the CPU cache on the machines I own.