Scaling-laws for Large Time-series Models
Edwards, Thomas D. P., Alvey, James, Alsing, Justin, Nguyen, Nam H., Wandelt, Benjamin D.
–arXiv.org Artificial Intelligence
Scaling laws for large language models (LLMs) have provided useful guidance on how to train ever larger models for predictable performance gains. Time series forecasting shares a similar sequential structure to language, and is amenable to large-scale transformer architectures. Here we show that foundational decoder-only time series transformer models exhibit analogous scaling-behavior to LLMs, while architectural details (aspect ratio and number of heads) have a minimal effect over broad ranges. We assemble a large corpus of heterogenous time series data on which to train, and establish, for the first time, power-law scaling relations with respect to parameter count, dataset size, and training compute, spanning five orders of magnitude.
arXiv.org Artificial Intelligence
May-22-2024
- Country:
- North America
- United States
- California (0.04)
- New York (0.04)
- Trinidad and Tobago > Trinidad
- United States
- Europe > Netherlands
- North Holland > Amsterdam (0.04)
- North America
- Genre:
- Research Report (1.00)
- Industry:
- Energy (1.00)
- Banking & Finance (0.93)
- Health & Medicine (0.68)
- Government > Regional Government (0.47)
- Technology: