"Paraphrasing The Original Text" Makes High Accuracy Long-Context QA

Yu, Yijiong

arXiv.org Artificial Intelligence 

However, currently no fine-tuning method on open-source datasets has achieved an LLM with Most open-source generative language models satisfactory long-context performance, and while refining the currently have a context window of no more than 4k, form of prompt may bring improvements to powerful LLMs limiting their ability when facing long text. Many [6], it may not work for those whose inherent long-context previous efforts have tried to extend the context ability is relatively weak. With this background, our research window of models, but their actual effects have focuses primarily on enhancing inherent long-context been found to be very limited. To address this issue, capabilities of LLM through lightweight fine-tuning on a we theoretically analyze the effectiveness of the low-cost constructed dataset, without specifically long-context training data and find that long-context designing the input format, significantly modifying the training requires "effective" data rather than simply model's structure, increasing its parameter scale or "long" data, which is rarely noticed in previous constructing expensive, labor-intensive and proprietary studies. Thus, we propose adding "original text datasets.