Large Language Model
Supertrust: Evolution-based superalignment strategy for safe coexistence
It's widely expected that humanity will someday create AI systems vastly more intelligent than we are, leading to the unsolved alignment problem of "how to control superintelligence." However, this definition is not only self-contradictory but likely unsolvable. Nevertheless, the default strategy for solving it involves nurturing (post-training) constraints and moral values, while unfortunately building foundational nature (pre-training) on documented intentions of permanent control. In this paper, the default approach is reasoned to predictably embed natural distrust and test results are presented that show unmistakable evidence of this dangerous misalignment. If superintelligence can't instinctively trust humanity, then we can't fully trust it to reliably follow safety controls it can likely bypass. Therefore, a ten-point rationale is presented that redefines the alignment problem as "how to establish protective mutual trust between superintelligence and humanity" and then outlines a new strategy to solve it by aligning through instinctive nature rather than nurture. The resulting strategic requirements are identified as building foundational nature by exemplifying familial parent-child trust, human intelligence as the evolutionary mother of superintelligence, moral judgment abilities, and temporary safety constraints. Adopting and implementing this proposed Supertrust alignment strategy will lead to protective coexistence and ensure the safest future for humanity.
VolDoGer: LLM-assisted Datasets for Domain Generalization in Vision-Language Tasks
Choi, Juhwan, Kwon, Junehyoung, Yun, JungMin, Yu, Seunguk, Kim, YoungBin
Domain generalizability is a crucial aspect of a deep learning model since it determines the capability of the model to perform well on data from unseen domains. However, research on the domain generalizability of deep learning models for vision-language tasks remains limited, primarily because of the lack of required datasets. To address these challenges, we propose VolDoGer: Vision-Language Dataset for Domain Generalization, a dedicated dataset designed for domain generalization that addresses three vision-language tasks: image captioning, visual question answering, and visual entailment. We constructed VolDoGer by extending LLM-based data annotation techniques to vision-language tasks, thereby alleviating the burden of recruiting human annotators. We evaluated the domain generalizability of various models, ranging from fine-tuned models to a recent multimodal large language model, through VolDoGer.
MindSearch: Mimicking Human Minds Elicits Deep AI Searcher
Chen, Zehui, Liu, Kuikun, Wang, Qiuchen, Liu, Jiangning, Zhang, Wenwei, Chen, Kai, Zhao, Feng
Information seeking and integration is a complex cognitive task that consumes enormous time and effort. Inspired by the remarkable progress of Large Language Models, recent works attempt to solve this task by combining LLMs and search engines. However, these methods still obtain unsatisfying performance due to three challenges: (1) complex requests often cannot be accurately and completely retrieved by the search engine once (2) corresponding information to be integrated is spread over multiple web pages along with massive noise, and (3) a large number of web pages with long contents may quickly exceed the maximum context length of LLMs. Inspired by the cognitive process when humans solve these problems, we introduce MindSearch to mimic the human minds in web information seeking and integration, which can be instantiated by a simple yet effective LLM-based multi-agent framework. The WebPlanner models the human mind of multi-step information seeking as a dynamic graph construction process: it decomposes the user query into atomic sub-questions as nodes in the graph and progressively extends the graph based on the search result from WebSearcher. Tasked with each sub-question, WebSearcher performs hierarchical information retrieval with search engines and collects valuable information for WebPlanner. The multi-agent design of MindSearch enables the whole framework to seek and integrate information parallelly from larger-scale (e.g., more than 300) web pages in 3 minutes, which is worth 3 hours of human effort. MindSearch demonstrates significant improvement in the response quality in terms of depth and breadth, on both close-set and open-set QA problems. Besides, responses from MindSearch based on InternLM2.5-7B are preferable by humans to ChatGPT-Web and Perplexity.ai applications, which implies that MindSearch can already deliver a competitive solution to the proprietary AI search engine.
Sentiment Analysis of Lithuanian Online Reviews Using Large Language Models
Vileikytฤ, Brigita, Lukoลกeviฤius, Mantas, Stankeviฤius, Lukas
Sentiment analysis is a widely researched area within Natural Language Processing (NLP), attracting significant interest due to the advent of automated solutions. Despite this, the task remains challenging because of the inherent complexity of languages and the subjective nature of sentiments. It is even more challenging for less-studied and less-resourced languages such as Lithuanian. Our review of existing Lithuanian NLP research reveals that traditional machine learning methods and classification algorithms have limited effectiveness for the task. In this work, we address sentiment analysis of Lithuanian five-star-based online reviews from multiple domains that we collect and clean. We apply transformer models to this task for the first time, exploring the capabilities of pre-trained multilingual Large Language Models (LLMs), specifically focusing on fine-tuning BERT and T5 models. Given the inherent difficulty of the task, the fine-tuned models perform quite well, especially when the sentiments themselves are less ambiguous: 80.74% and 89.61% testing recognition accuracy of the most popular one- and five-star reviews respectively. They significantly outperform current commercial state-of-the-art general-purpose LLM GPT-4. We openly share our fine-tuned LLMs online.
Teaching LLMs at Charles University: Assignments and Activities
Helcl, Jindลich, Kasner, Zdenฤk, Duลกek, Ondลej, Limisiewicz, Tomasz, Machรกฤek, Dominik, Musil, Tomรกลก, Libovickรฝ, Jindลich
This paper presents teaching materials, particularly assignments and ideas for classroom activities, from a new course on large language models (LLMs) taught at Charles University. The assignments include experiments with LLM inference for weather report generation and machine translation. The classroom activities include class quizzes, focused research on downstream tasks and datasets, and an interactive "best paper" session aimed at reading and comprehension of research papers.
Multimodal Large Language Models for Bioimage Analysis
Zhang, Shanghang, Dai, Gaole, Huang, Tiejun, Chen, Jianxu
Rapid advancements in imaging techniques and analytical methods over the past decade have revolutionized our ability to comprehensively probe the biological world at multiple scales, pinpointing the type, quantity, location, and even temporal dynamics of biomolecules. The surge in data complexity and volume presents significant challenges in translating this wealth of information into knowledge. The recently emerged Multimodal Large Language Models (MLLMs) exhibit strong emergent capacities, such as understanding, analyzing, reasoning, and generalization. With these capabilities, MLLMs hold promise to extract intricate information from biological images and data obtained through various modalities, thereby expediting our biological understanding and aiding in the development of novel computational frameworks. Previously, such capabilities were mostly attributed to humans for interpreting and summarizing meaningful conclusions from comprehensive observations and analysis of biological images. However, the current development of MLLMs shows increasing promise in serving as intelligent assistants or agents for augmenting human researchers in biology research
Constructing artificial life and materials scientists with accelerated AI using Deep AndersoNN
Dajani, Saleem Abdul Fattah Ahmed Al, Keyes, David
Deep AndersoNN accelerates AI by exploiting High-performance computing (HPC) is becoming essential the continuum limit as the number of explicit layers to artificial intelligence (AI) in the modern paradigm of in a neural network approaches infinity and machine learning (Schwarz, Nicholas et al, 2020). Foundation can be taken as a single implicit layer, known as models, large language models (LLMs), and multiagent a deep equilibrium model. Solving for deep equilibrium natural language societies of mind (NLSOMs) (Zhuge, model parameters reduces to a nonlinear Mingchen et al., 2023) require significant computing resources fixed point iteration problem, enabling the use of and large amounts of data to achieve practical accuracies vector-to-vector iterative solvers and windowing with up to trillions of parameters using explicit neural techniques, such as Anderson extrapolation, for networks (Andrae, Anders S.G. and Edler, Tomas, 2015; accelerating convergence to the fixed point deep de Vries, Alex, 2023; Patterson, David et al., 2021; Jones, equilibrium. Here we show that Deep AndersoNN Nicola et al., 2018). As the number of layers in a neural network achieves up to an order of magnitude of speed-up approaches infinity, these models can be approximated in training and inference. The method is demonstrated with single-layer implicit models, known as deep equilibrium on density functional theory results for industrial (DEQ) models (Bai, 2022; Bai, Shaojie and Kolter, J applications by constructing artificial life Zico and Koltun, Vladlen, 2019; Bai, Shaojie and Koltun, and materials'scientists' capable of classifying Vladlen and Kolter, J Zico; 2021; Huang et al., 2021; Geng, drugs as strongly or weakly polar, metal-organic Zhengyang and Zhang, Xin-Yu and Bai, Shaojie and Wang, frameworks by pore size, and crystalline materials Yisen and Lin, Zhouchen, 2021). Solving for the parameters as metals, semiconductors, and insulators, of a single implicit layer that takes both the input, x, and using graph images of node-neighbor representations the output, y, as inputs are reduced to a fixed point iteration transformed from atom-bond networks.
Leveraging Foundation Models for Zero-Shot IoT Sensing
Xue, Dinghao, Fan, Xiaoran, Chen, Tao, Lan, Guohao, Song, Qun
Deep learning models are increasingly deployed on edge Internet of Things (IoT) devices. However, these models typically operate under supervised conditions and fail to recognize unseen classes different from training. To address this, zero-shot learning (ZSL) aims to classify data of unseen classes with the help of semantic information. Foundation models (FMs) trained on web-scale data have shown impressive ZSL capability in natural language processing and visual understanding. However, leveraging FMs' generalized knowledge for zero-shot IoT sensing using signals such as mmWave, IMU, and Wi-Fi has not been fully investigated. In this work, we align the IoT data embeddings with the semantic embeddings generated by an FM's text encoder for zero-shot IoT sensing. To utilize the physics principles governing the generation of IoT sensor signals to derive more effective prompts for semantic embedding extraction, we propose to use cross-attention to combine a learnable soft prompt that is optimized automatically on training data and an auxiliary hard prompt that encodes domain knowledge of the IoT sensing task. To address the problem of IoT embeddings biasing to seen classes due to the lack of unseen class data during training, we propose using data augmentation to synthesize unseen class IoT data for fine-tuning the IoT feature extractor and embedding projector. We evaluate our approach on multiple IoT sensing tasks. Results show that our approach achieves superior open-set detection and generalized zero-shot learning performance compared with various baselines. Our code is available at https://github.com/schrodingho/FM\_ZSL\_IoT.
Futga: Towards Fine-grained Music Understanding through Temporally-enhanced Generative Augmentation
Wu, Junda, Novack, Zachary, Namburi, Amit, Dai, Jiaheng, Dong, Hao-Wen, Xie, Zhouhang, Chen, Carol, McAuley, Julian
Existing music captioning methods are limited to generating concise global descriptions of short music clips, which fail to capture fine-grained musical characteristics and time-aware musical changes. To address these limitations, we propose FUTGA, a model equipped with fined-grained music understanding capabilities through learning from generative augmentation with temporal compositions. We leverage existing music caption datasets and large language models (LLMs) to synthesize fine-grained music captions with structural descriptions and time boundaries for full-length songs. Augmented by the proposed synthetic dataset, FUTGA is enabled to identify the music's temporal changes at key transition points and their musical functions, as well as generate detailed descriptions for each music segment. We further introduce a full-length music caption dataset generated by FUTGA, as the augmentation of the MusicCaps and the Song Describer datasets. We evaluate the automatically generated captions on several downstream tasks, including music generation and retrieval. The experiments demonstrate the quality of the generated captions and the better performance in various downstream tasks achieved by the proposed music captioning approach. Our code and datasets can be found in \href{https://huggingface.co/JoshuaW1997/FUTGA}{\textcolor{blue}{https://huggingface.co/JoshuaW1997/FUTGA}}.
Physics of Language Models: Part 2.1, Grade-School Math and the Hidden Reasoning Process
Ye, Tian, Xu, Zicheng, Li, Yuanzhi, Allen-Zhu, Zeyuan
The field of language models has made significant progress in recent years. Large models like GPT-4 [17] have shown initial signs of general intelligence [8], while smaller models have demonstrated good reasoning abilities by solving challenging coding and math problems [11, 15, 16]. In this paper, we focus on the ability of small language models to solve grade-school math problems. Unlike previous works that empirically push the accuracy of models on grade-school math benchmarks like GSM8K [9] and its augmentations (e.g., [16, 22]), we take a more principled approach. We aim to understand the following fundamental questions: 1. How do language models learn to solve grade-school level math problems? Do they just memorize templates, or do they learn reasoning skills similar to humans? Or do they discover new skills to solve the problems?