HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs

Open in new window