Revisiting Intermediate Layer Distillation for Compressing Language Models: An Overfitting Perspective

Open in new window