Deep Learning
Memory-Efficient Fine-Tuning of Compressed Large Language Models via sub-4-bit Integer Quantization
While parameter-efficient fine-tuning (PEFT) methods aim to reduce the memory usage of the optimizer state during fine-tuning, the inherent size of pre-trained LLM weights continues to be a pressing concern. Even though quantization techniques are widely proposed to ease memory demands and accelerate LLM inference, most of these techniques are geared towards the deployment phase.
Reviews: A Primal Dual Formulation For Deep Learning With Constraints
NeurIPS 2019 Sun Dec 8th through Sat the 14th, 2019 at Vancouver Convention Center "6594" "A Primal Dual Formulation For Deep Learning With Constraints" All reviewers were positive about the contributions in the paper so I recommend acceptance. Please take into account all the reviewers' comments when preparing the final version of the paper.