CSKV: Training-Efficient Channel Shrinking for KV Cache in Long-Context Scenarios

Open in new window