Large Language Model
Towards Effective and Efficient Continual Pre-training of Large Language Models
Chen, Jie, Chen, Zhipeng, Wang, Jiapeng, Zhou, Kun, Zhu, Yutao, Jiang, Jinhao, Min, Yingqian, Zhao, Wayne Xin, Dou, Zhicheng, Mao, Jiaxin, Lin, Yankai, Song, Ruihua, Xu, Jun, Chen, Xu, Yan, Rui, Wei, Zhewei, Hu, Di, Huang, Wenbing, Wen, Ji-Rong
Continual pre-training (CPT) has been an important approach for adapting language models to specific domains or tasks. To make the CPT approach more traceable, this paper presents a technical report for continually pre-training Llama-3 (8B), which significantly enhances the Chinese language ability and scientific reasoning ability of the backbone model. To enhance the new abilities while retaining the original abilities, we design specific data mixture and curriculum strategies by utilizing existing datasets and synthesizing high-quality datasets. Specifically, we synthesize multidisciplinary scientific question and answer (QA) pairs based on related web pages, and subsequently incorporate these synthetic data to improve the scientific reasoning ability of Llama-3. We refer to the model after CPT as Llama-3-SynE (Synthetic data Enhanced Llama-3). We also present the tuning experiments with a relatively small model -- TinyLlama, and employ the derived findings to train the backbone model. Extensive experiments on a number of evaluation benchmarks show that our approach can largely improve the performance of the backbone models, including both the general abilities (+8.81 on C-Eval and +6.31 on CMMLU) and the scientific reasoning abilities (+12.00 on MATH and +4.13 on SciEval), without hurting the original capacities. Our model, data, and codes are available at https://github.com/RUC-GSAI/Llama-3-SynE.
Greedy Output Approximation: Towards Efficient Structured Pruning for LLMs Without Retraining
Li, Jianwei, Dong, Yijun, Lei, Qi
To remove redundant components of large language models (LLMs) without incurring significant computational costs, this work focuses on single-shot pruning without a retraining phase. We simplify the pruning process for Transformer-based LLMs by identifying a depth-2 pruning structure that functions independently. Additionally, we propose two inference-aware pruning criteria derived from the optimization perspective of output approximation, which outperforms traditional training-aware metrics such as gradient and Hessian. We also introduce a two-step reconstruction technique to mitigate pruning errors without model retraining. Experimental results demonstrate that our approach significantly reduces computational costs and hardware requirements while maintaining superior performance across various datasets and models.
Adaptive Contrastive Search: Uncertainty-Guided Decoding for Open-Ended Text Generation
Arias, Esteban Garces, Rodemann, Julian, Li, Meimingwei, Heumann, Christian, Aßenmacher, Matthias
Decoding from the output distributions of large language models to produce high-quality text is a complex challenge in language modeling. Various approaches, such as beam search, sampling with temperature, $k-$sampling, nucleus $p-$sampling, typical decoding, contrastive decoding, and contrastive search, have been proposed to address this problem, aiming to improve coherence, diversity, as well as resemblance to human-generated text. In this study, we introduce adaptive contrastive search, a novel decoding strategy extending contrastive search by incorporating an adaptive degeneration penalty, guided by the estimated uncertainty of the model at each generation step. This strategy is designed to enhance both the creativity and diversity of the language modeling process while at the same time producing coherent and high-quality generated text output. Our findings indicate performance enhancement in both aspects, across different model architectures and datasets, underscoring the effectiveness of our method in text generation tasks. Our code base, datasets, and models are publicly available.
Cluster-norm for Unsupervised Probing of Knowledge
Laurito, Walter, Maiya, Sharan, Dhimoïla, Grégoire, Owen, null, Yeung, null, Hänni, Kaarel
The deployment of language models brings challenges in generating reliable information, especially when these models are fine-tuned using human preferences. To extract encoded knowledge without (potentially) biased human labels, unsupervised probing techniques like Contrast-Consistent Search (CCS) have been developed (Burns et al., 2022). However, salient but unrelated features in a given dataset can mislead these probes (Farquhar et al., 2023). Addressing this, we propose a cluster normalization method to minimize the impact of such features by clustering and normalizing activations of contrast pairs before applying unsupervised probing techniques. While this approach does not address the issue of differentiating between knowledge in general and simulated knowledge - a major issue in the literature of latent knowledge elicitation (Christiano et al., 2021) - it significantly improves the ability of unsupervised probes to identify the intended knowledge amidst distractions.
Blockchain for Large Language Model Security and Safety: A Holistic Survey
Geren, Caleb, Board, Amanda, Dagher, Gaby G., Andersen, Tim, Zhuang, Jun
With the advent of accessible interfaces for interacting with large language models, there has been an associated explosion in both their commercial and academic interest. Consequently, there has also been an sudden burst of novel attacks associated with large language models, jeopardizing user data on a massive scale. Situated at a comparable crossroads in its development, and equally prolific to LLMs in its rampant growth, blockchain has emerged in recent years as a disruptive technology with the potential to redefine how we approach data handling. In particular, and due to its strong guarantees about data immutability and irrefutability as well as inherent data provenance assurances, blockchain has attracted significant attention as a means to better defend against the array of attacks affecting LLMs and further improve the quality of their responses. In this survey, we holistically evaluate current research on how blockchains are being used to help protect against LLM vulnerabilities, as well as analyze how they may further be used in novel applications. To better serve these ends, we introduce a taxonomy of blockchain for large language models (BC4LLM) and also develop various definitions to precisely capture the nature of different bodies of research in these areas. Moreover, throughout the paper, we present frameworks to contextualize broader research efforts, and in order to motivate the field further, we identify future research goals as well as challenges present in the blockchain for large language model (BC4LLM) space.
Climbing the Complexity Ladder with Expressive Attention
Attention involves comparing query and key vectors in terms of a scalar product, $\mathbf{Q}^T\mathbf{K}$, together with a subsequent softmax normalization. Classicaly, parallel/orthogonal/antiparallel queries and keys lead to large/intermediate/small attention weights. Here we study expressive attention (EA), which is based on $(\mathbf{Q}^T\mathbf{K})^2$, the squared dot product. In this case attention is enhanced when query and key are either parallel or antiparallel, and suppressed for orthogonal configurations. For a series of autoregressive prediction tasks, we find that EA performs at least as well as the standard mechanism, dot-product attention (DPA). Increasing task complexity, EA is observed to outperform DPA with increasing margins, which also holds for multi-task settings. For a given model size, EA manages to achieve 100\% performance for a range of complexity levels not accessible to DPA.
OopsGPT
Whenever AI companies present a vision for the role of artificial intelligence in the future of searching the internet, they tend to underscore the same points: instantaneous summaries of relevant information; ready-made lists tailored to a searcher's needs. They tend not to point out that generative-AI models are prone to providing incorrect, and at times fully made-up, information--and yet it keeps happening. Early this afternoon, OpenAI, the maker of ChatGPT, announced a prototype AI tool that can search the web and answer questions, fittingly called SearchGPT. The launch is designed to hint at how AI will transform the ways in which people navigate the internet--except that, before users have had a chance to test the new program, it already appears error prone. The tool then pulls up a list of festivals that it states are taking place in Boone this August, the first being An Appalachian Summer Festival, which according to the tool is hosting a series of arts events from July 29 to August 16 of this year. Someone in Boone hoping to buy tickets to one of those concerts, however, would run into trouble.
SearchGPT Is OpenAI's Direct Assault on Google
After months of speculation about its search ambitions, OpenAI has revealed SearchGPT, a "prototype" search engine that could eventually help the company tear off a slice of Google's lucrative business. OpenAI said that the new tool would help users find what they are looking for more quickly and easily by using generative AI to gather links and answer user queries in a conversational tone. In addition to a broader web search, the search engine will tap into information provided by publishers who have signed deals giving OpenAI access to their data. Kayla Wood, a spokesperson for OpenAI, declined to provide a SearchGPT demo or an interview about the new tool for WIRED, but confirmed that the company has already given access to unnamed partners and publishers and improved aspects of the search engine based on their feedback. Microsoft, an investor in OpenAI, was one of the first companies to release a generative AI search engine to the public when it launched an AI-powered version of Bing back in 2023 that relied on OpenAI's large language models.
OpenAI unveils SearchGPT, an AI-powered search engine
OpenAI on Thursday announced a new AI-powered search engine prototype called SearchGPT. The move marks the company's entry into a competitive search engine market dominated by Google for decades. On its website, OpenAI described SearchGPT as "a temporary prototype of new AI search features that give you fast and timely answers with clear and relevant sources." The company plans to test out the product with 10,000 initial users and then roll it into ChatGPT after gathering feedback. The launch of SearchGPT comes amid growing competition in AI-powered search.
ChatGPT has its own AI search engine now
In order to train their models, AI generative text tools like ChatGPT scour the internet for text…which is also something that search engines like Google do. So, why not combine them and just give you everything? That seems to be the thinking behind SearchGPT, a new search engine from ChatGPT maker OpenAI. The product was announced as a prototype on OpenAI's website, inviting users to join a wait list to access the tool. According to the company, it's designed to "combine the strength of our AI models with information from the web to give you fast and timely answers with clear and relevant resources."