Dual-Head Reasoning Distillation: Improving Classifier Accuracy with Train-Time-Only Reasoning

Open in new window