LoRA Training in the NTK Regime has No Spurious Local Minima
Jang, Uijeong, Lee, Jason D., Ryu, Ernest K.
–arXiv.org Artificial Intelligence
Low-rank adaptation (LoRA) has become the standard approach for parameter-efficient fine-tuning of large language models (LLM), but our theoretical understanding of LoRA has been limited. In this work, we theoretically analyze LoRA fine-tuning in the neural tangent kernel (NTK) regime with $N$ data points, showing: (i) full fine-tuning (without LoRA) admits a low-rank solution of rank $r\lesssim \sqrt{N}$; (ii) using LoRA with rank $r\gtrsim \sqrt{N}$ eliminates spurious local minima, allowing gradient descent to find the low-rank solutions; (iii) the low-rank solution found using LoRA generalizes well.
arXiv.org Artificial Intelligence
May-28-2024
- Country:
- North America > United States
- New York (0.04)
- California > Los Angeles County
- Los Angeles (0.14)
- Europe > Austria
- Vienna (0.14)
- Asia
- Middle East > Jordan (0.04)
- South Korea > Seoul
- Seoul (0.04)
- North America > United States
- Genre:
- Research Report (0.64)
- Technology: