mofo
MoFo: Empowering Long-term Time Series Forecasting with Periodic Pattern Modeling
The stable periodic patterns present in the time series data serve as the foundation for long-term forecasting. However, existing models suffer from limitations such as continuous and chaotic input partitioning, as well as weak inductive biases, which restrict their ability to capture such recurring structures. In this paper, we propose MoFo, which interprets periodicity as both the correlation of periodaligned time steps and the trend of period-offset time steps. We first design periodstructured patches--2D tensors generated through discrete sampling--where each row contains only period-aligned time steps, enabling direct modeling of periodic correlations. Period-offset time steps within a period are aligned in columns.
MoFO: Momentum-Filtered Optimizer for Mitigating Forgetting in LLM Fine-Tuning
Chen, Yupeng, Wang, Senmiao, Lin, Zhihang, Qin, Zeyu, Zhang, Yushun, Ding, Tian, Sun, Ruoyu
Recently, large language models (LLMs) have demonstrated remarkable capabilities in a wide range of tasks. Typically, an LLM is pre-trained on large corpora and subsequently fine-tuned on task-specific datasets. However, during fine-tuning, LLMs may forget the knowledge acquired in the pre-training stage, leading to a decline in general capabilities. To address this issue, we propose a new fine-tuning algorithm termed Momentum-Filtered Optimizer (MoFO). The key idea of MoFO is to iteratively select and update the model parameters with the largest momentum magnitudes. Compared to full-parameter training, MoFO achieves similar fine-tuning performance while keeping parameters closer to the pre-trained model, thereby mitigating knowledge forgetting. Unlike most existing methods for forgetting mitigation, MoFO combines the following two advantages. First, MoFO does not require access to pre-training data. This makes MoFO particularly suitable for fine-tuning scenarios where pre-training data is unavailable, such as fine-tuning checkpoint-only open-source LLMs. Second, MoFO does not alter the original loss function. This could avoid impairing the model performance on the fine-tuning tasks. We validate MoFO through rigorous convergence analysis and extensive experiments, demonstrating its superiority over existing methods in mitigating forgetting and enhancing fine-tuning performance.
Judge in Uber-Waymo suit says Google co-founder Sergey Brin 'better show up'
In the latest hearing in the Uber vs. Waymo lawsuit on Wednesday, San Francisco district judge William Alsup addressed Uber's complaint that Google co-founder Sergey Brin is trying to avoid deposition. Alsup said, "you go back and tell that guy he better show up," after voicing frustration at Alphabet executives claiming they are "too busy." Brin is currently the president of Alphabet, the holding company that includes both Google and Waymo, the self-driving car unit that was spun out of Google. Alsup also said Anthony Levandowski, the engineer at the center of the dispute, could be called to testify in court even though he has previously pleaded the fifth. Alsup also said that all questions by Uber and Waymo would be reviewed in advance and must have evidence backing them.