TOKON: TOKenization-Optimized Normalization for time series analysis with a large language model

Yang, Janghoon

arXiv.org Artificial Intelligence 

-- While large language models have rapidly evolved towards general artificial intelligence, their versatility in analyzing time series data remains limited. To address this limitation, we propose a novel normalization technique that considers the inherent nature of tokenization. The proposed Tokenization-Optimized Normalization (TOKON) simplifies time series data by representing each element with a single token, effectively reducing the number of tokens by 2 to 3 times. Additionally, we introduce a novel prompt for time series forecasting, termed Time Series Forecasting with Care (TFSC), to further enhance forecasting performance. Experimental results demonstrate that TOKON improves root mean square error (RMSE) for multi-step forecasting by approximately 7% to 18%, depending on the dataset and prompting method. With the evolution of deep learning in natural language processing, the ubiquitous nature of large language models (LLMs) is becoming increasingly robust [1]. Initially introduced for Q&A services, these models can now be applied to sound and image modalities, enabling LLMs to understand and integrate multi-modal information [2]. The extensive scale of these models allows them to achieve state-of-the-art (SOTA) performance across various tasks, including natural language understanding, event recognition, and coding.

Duplicate Docs Excel Report

Title
None found

Similar Docs  Excel Report  more

TitleSimilaritySource
None found