Reducing Transformer Key-Value Cache Size with Cross-Layer Attention William Brandon

Open in new window