Greedy Output Approximation: Towards Efficient Structured Pruning for LLMs Without Retraining

Open in new window