cepe
Long-Context Language Modeling with Parallel Context Encoding
Yen, Howard, Gao, Tianyu, Chen, Danqi
Extending large language models (LLMs) to process longer inputs is crucial for a wide range of applications. However, the substantial computational cost of transformers and limited generalization of positional encoding restrict the size of their context window. We introduce Context Expansion with Parallel Encoding (CEPE), a framework that can be applied to any existing decoder-only LLMs to extend their context window. CEPE employs a small encoder to process long inputs chunk by chunk, enabling the frozen decoder to utilize additional contexts via cross-attention. CEPE is efficient, generalizable, and versatile: trained with 8K-token documents, it extends the context window of LLAMA-2 to 128K tokens, offering 10x the throughput with only 1/6 of the memory. CEPE yields strong performance on language modeling and in-context learning. CEPE also excels in retrieval-augmented applications, while existing long-context models degenerate with retrieved contexts. We further introduce a CEPE variant that can extend the context window of instruction-tuned models using only unlabeled data, and showcase its effectiveness on LLAMA-2-CHAT, leading to a strong instruction-following model that can leverage very long contexts on downstream tasks.
Collaborative Learning in Kernel-based Bandits for Distributed Users
Salgia, Sudeep, Vakili, Sattar, Zhao, Qing
We study collaborative learning among distributed clients facilitated by a central server. Each client is interested in maximizing a personalized objective function that is a weighted sum of its local objective and a global objective. Each client has direct access to random bandit feedback on its local objective, but only has a partial view of the global objective and relies on information exchange with other clients for collaborative learning. We adopt the kernel-based bandit framework where the objective functions belong to a reproducing kernel Hilbert space. We propose an algorithm based on surrogate Gaussian process (GP) models and establish its order-optimal regret performance (up to polylogarithmic factors). We also show that the sparse approximations of the GP models can be employed to reduce the communication overhead across clients.
American University of Sharjah launches Certificate in Artificial Intelligence for Smart Cities
With the creation of Smart Cities high on the UAE government's agenda, a new course from the Center for Executive and Professional Education (CEPE) at American University of Sharjah (AUS) will help executives apply the benefits of Artificial Intelligence (AI) to their Smart City projects. The Certificate in AI for Smart Cities is being delivered by CEPE in conjunction with the AUS College of Engineering's Department of Computer Science and Engineering. Topics to be covered include Big Data, Machine Learning, Cyber-Physical Systems, Internet of Things, Cyber Security and Blockchain, and Cloud Computing. Participants require no prior knowledge of AI, as the course has been designed for mid- to high-level executives from across the Middle East tasked with developing Smart City solutions. It is an introductory course intended for those who want to better understand how AI can transform their operations, with a focus on learning from global best practice and case studies. AI holds enormous value for the UAE and wider GCC.