Large Language Model
CUTE: Measuring LLMs' Understanding of Their Tokens
Edman, Lukas, Schmid, Helmut, Fraser, Alexander
Large Language Models (LLMs) show remarkable performance on a wide variety of tasks. Most LLMs split text into multi-character tokens and process them as atomic units without direct access to individual characters. This raises the question: To what extent can LLMs learn orthographic information? To answer this, we propose a new benchmark, CUTE, which features a collection of tasks designed to test the orthographic knowledge of LLMs. We evaluate popular LLMs on CUTE, finding that most of them seem to know the spelling of their tokens, yet fail to use this information effectively to manipulate text, calling into question how much of this knowledge is generalizable.
Reliable and diverse evaluation of LLM medical knowledge mastery
Zhou, Yuxuan, Liu, Xien, Ning, Chen, Zhang, Xiao, Wu, Ji
Mastering medical knowledge is crucial for medical-specific LLMs. However, despite the existence of medical benchmarks like MedQA, a unified framework that fully leverages existing knowledge bases to evaluate LLMs' mastery of medical knowledge is still lacking. In the study, we propose a novel framework PretexEval that dynamically generates reliable and diverse test samples to evaluate LLMs for any given medical knowledge base. We notice that test samples produced directly from knowledge bases by templates or LLMs may introduce factual errors and also lack diversity. To address these issues, we introduce a novel schema into our proposed evaluation framework that employs predicate equivalence transformations to produce a series of variants for any given medical knowledge point. Finally, these produced predicate variants are converted into textual language, resulting in a series of reliable and diverse test samples to evaluate whether LLMs fully master the given medical factual knowledge point. Here, we use our proposed framework to systematically investigate the mastery of medical factual knowledge of 12 well-known LLMs, based on two knowledge bases that are crucial for clinical diagnosis and treatment. The evaluation results illustrate that current LLMs still exhibit significant deficiencies in fully mastering medical knowledge, despite achieving considerable success on some famous public benchmarks. These new findings provide valuable insights for developing medical-specific LLMs, highlighting that current LLMs urgently need to strengthen their comprehensive and in-depth mastery of medical knowledge before being applied to real-world medical scenarios.
Advancing Event Causality Identification via Heuristic Semantic Dependency Inquiry Network
Li, Haoran, Gao, Qiang, Wu, Hongmei, Huang, Li
Event Causality Identification (ECI) focuses on extracting causal relations between events in texts. Existing methods for ECI primarily rely on causal features and external knowledge. However, these approaches fall short in two dimensions: (1) causal features between events in a text often lack explicit clues, and (2) external knowledge may introduce bias, while specific problems require tailored analyses. To address these issues, we propose SemDI - a simple and effective Semantic Dependency Inquiry Network for ECI. SemDI captures semantic dependencies within the context using a unified encoder. Then, it utilizes a Cloze Analyzer to generate a fill-in token based on comprehensive context understanding. Finally, this fill-in token is used to inquire about the causal relation between two events. Extensive experiments demonstrate the effectiveness of SemDI, surpassing state-of-the-art methods on three widely used benchmarks. Code is available at https://github.com/hrlics/SemDI.
Contextual Compression in Retrieval-Augmented Generation for Large Language Models: A Survey
Large Language Models (LLMs) showcase remarkable abilities, yet they struggle with limitations such as hallucinations, outdated knowledge, opacity, and inexplicable reasoning. To address these challenges, Retrieval-Augmented Generation (RAG) has proven to be a viable solution, leveraging external databases to improve the consistency and coherence of generated content, especially valuable for complex, knowledge-rich tasks, and facilitates continuous improvement by leveraging domain-specific insights. By combining the intrinsic knowledge of LLMs with the vast, dynamic repositories of external databases, RAG achieves a synergistic effect. However, RAG is not without its limitations, including a limited context window, irrelevant information, and the high processing overhead for extensive contextual data. In this comprehensive work, we explore the evolution of Contextual Compression paradigms, providing an in-depth examination of the field. Finally, we outline the current challenges and suggest potential research and development directions, paving the way for future advancements in this area.
Anchors Aweigh! Sail for Optimal Unified Multi-Modal Representations
Jeong, Minoh, Namgung, Min, Kim, Zae Myung, Kang, Dongyeop, Chiang, Yao-Yi, Hero, Alfred
Multimodal learning plays a crucial role in enabling machine learning models to fuse and utilize diverse data sources, such as text, images, and audio, to support a variety of downstream tasks. A unified representation across various modalities is particularly important for improving efficiency and performance. Recent binding methods, such as ImageBind (Girdhar et al., 2023), typically use a fixed anchor modality to align multimodal data in the anchor modal embedding space. In this paper, we mathematically analyze the fixed anchor binding methods and uncover notable limitations: (1) over-reliance on the choice of the anchor modality, (2) failure to capture intra-modal information, and (3) failure to account for inter-modal correlation among non-anchored modalities. To address these limitations, we propose CentroBind, a simple yet powerful approach that eliminates the need for a fixed anchor; instead, it employs dynamically adjustable centroid-based anchors generated from all available modalities, resulting in a balanced and rich representation space. We theoretically demonstrate that our method captures three crucial properties of multimodal learning: intra-modal learning, inter-modal learning, and multimodal alignment, while also constructing a robust unified representation across all modalities. Our experiments on both synthetic and real-world datasets demonstrate the superiority of the proposed method, showing that dynamic anchor methods outperform all fixed anchor binding methods as the former captures more nuanced multimodal interactions.
Microsoft's Copilot AI gets a voice and the ability to see websites you browse
Beyond debuting new features for Copilot AI PCs and Windows 11's 2024 update, Microsoft is also giving its Copilot AI a makeover on the web, mobile and desktop. That includes a slightly friendlier interface wherever you access it, along with new capabilities like Copilot Voice, which allows you to talk conversationally with the AI assistant. Ultimately, Microsoft is aiming for Copilot to be seen as more than just a party trick for generative AI search and image creation -- it's trying to make it a core part of your daily workflow. That starts with a cleaner and simpler UI that makes Copilot look different than a boring old search engine. You'll also be able to access Copilot from within Whatsapp, which could be useful if you want to avoid Meta's AI assistant.
Microsoft's AI Boss Wants Copilot to Bring 'Emotional Support' to Windows and Office
Mustafa Suleyman was at the center of an artificial intelligence revolution once before. As a cofounder of DeepMind, a British company acquired by Google in 2014, he helped devise a new way for computers to tackle seemingly impossible problems by combining practice with positive and negative feedback. DeepMind demonstrated the approach by developing a superhuman Go-playing program, AlphaGo, which defeated the world's best Go player in 2016. Now the CEO of Microsoft AI, Suleyman is talking up a new kind of AI breakthrough. As CEO of Microsoft AI, Suleyman oversees efforts to integrate the same AI that powers ChatGPT into software--including the Windows operating system--that runs most of the world's personal computers.
Copilot's AI will be able to 'see' and talk to you, Microsoft says
Microsoft is beginning to roll out its next feature update of Windows 11, the Windows 11 2024 Update, beginning today. But Microsoft obviously isn't done yet, and it's offering a sneak peek at new Copilot experiences which will debut this fall, including Copilot Voice, Copilot Vision, and Copilot Daily, among others. On the surface, the new additions to Copilot sound similar to multimodal ChatGPT (or GPT-4o) that OpenAI launched earlier this year, where ChatGPT can now "see" and an Advanced Voice feature means that you can have conversations with it. But there are some key differences between what Microsoft and OpenAI are offering, and only some of Microsoft's Copilot innovations will be available right away. It's probably safe to say, though, that Copilot Voice will be the most important addition -- and Copilot Vision may not be.
Microsoft's Copilot AI Gets a Voice, Vision, and a 'Hype Man' Persona
Microsoft deleted the over-eager office assistant Clippy some 17 years ago, but the vision for an friendly and optimistic AI helper has apparently found its way out of the Recycle Bin. The company is overhauling Copilot, the text-based artificial intelligence tool bundled with Windows and other software, with the addition of vision, voice, and the ability to solve more complex problems--along with a more "encouraging" personality. "We really are at this amazing kind of transition point," says Mustafa Suleyman, CEO of Microsoft AI. "AI companions now see what we see, hear what we hear, and speak in the same language that we use to communicate with one another." Copilot has so far met with a mixed response, with some users complaining of lag or vagueness in its responses, but Microsoft is betting that the tool could eventually become an integral part of Windows, Office, and beyond. By incorporating OpenAI's AI algorithms into software that is used by hundreds of millions of people, the company is also at the forefront of testing the potential for AI to boost productivity in office work.
On-Device Collaborative Language Modeling via a Mixture of Generalists and Specialists
Fan, Dongyang, Messmer, Bettina, Jaggi, Martin
On-device LLMs have gained increasing attention for their ability to enhance privacy and provide a personalized user experience. To facilitate learning with private and scarce local data, federated learning has become a standard approach, though it introduces challenges related to system and data heterogeneity among end users. As a solution, we propose a novel Collaborative learning approach with a Mixture of Generalists and Specialists (CoMiGS), being the first to effectively address both. Our approach distinguishes generalists and specialists by aggregating certain experts across end users while keeping others localized to specialize in user-specific datasets. A key innovation of our method is the bi-level optimization formulation of the Mixture-of-Experts learning objective, where the router is updated using a separate validation set that represents the target distribution. CoMiGS effectively balances collaboration and personalization, as demonstrated by its superior performance in scenarios with high data heterogeneity across multiple datasets. By decoupling resource abundance from data quantity, CoMiGS remains robust against overfitting--due to the generalists' regularizing effect--while adapting to local data through specialist expertise. Large Language Models (LLMs) have been showing great success serving as foundation models, evidenced by their capability to understand a wide range of tasks, such as ChatGPT (OpenAI, 2023), Claude (Anthropic, 2023), Gemini (DeepMind, 2023) and etc. However, cloud-based inference introduces significant delays for end users, and it often fails to meet their personalized needs (Ding et al., 2024; Iyengar & Adusumilli, 2024). Recently, there has been growing interest in deploying LLMs on edge devices, which offer benefits like lower latency, data localization, and more personalized user experiences (Xu et al., 2024). For instance, Apple (2024) recently launched on-device foundation models as part of its personal intelligence system. On-device LLMs present challenges such as limited and variable computational resources, scarce and heterogeneous local data, and privacy concerns related to data sharing (Peng et al., 2024; Wagner et al., 2024).