Large Language Model
Large Language Models (LLMs) as Agents for Augmented Democracy
Gudiño-Rosero, Jairo, Grandi, Umberto, Hidalgo, César A.
We explore the capabilities of an augmented democracy system built on off-the-shelf LLMs fine-tuned on data summarizing individual preferences across 67 policy proposals collected during the 2022 Brazilian presidential elections. We use a train-test cross-validation setup to estimate the accuracy with which the LLMs predict both: a subject's individual political choices and the aggregate preferences of the full sample of participants. At the individual level, the accuracy of the out of sample predictions lie in the range 69%-76% and are significantly better at predicting the preferences of liberal and college educated participants. At the population level, we aggregate preferences using an adaptation of the Borda score and compare the ranking of policy proposals obtained from a probabilistic sample of participants and from data augmented using LLMs. We find that the augmented data predicts the preferences of the full population of participants better than probabilistic samples alone when these represent less than 30% to 40% of the total population. These results indicate that LLMs are potentially useful for the construction of systems of augmented democracy.
An LLM-Tool Compiler for Fused Parallel Function Calling
Singh, Simranjit, Karatzas, Andreas, Fore, Michael, Anagnostopoulos, Iraklis, Stamoulis, Dimitrios
State-of-the-art sequential reasoning in Large Language Models (LLMs) has expanded the capabilities of Copilots beyond conversational tasks to complex function calling, managing thousands of API calls. However, the tendency of compositional prompting to segment tasks into multiple steps, each requiring a round-trip to the GPT APIs, leads to increased system latency and costs. Although recent advancements in parallel function calling have improved tool execution per API call, they may necessitate more detailed in-context instructions and task breakdown at the prompt level, resulting in higher engineering and production costs. Inspired by the hardware design principles of multiply-add (MAD) operations, which fuse multiple arithmetic operations into a single task from the compiler's perspective, we propose LLM-Tool Compiler, which selectively fuses similar types of tool operations under a single function at runtime, presenting them as a unified task to the LLM. This selective fusion inherently enhances parallelization and efficiency. Benchmarked on a large-scale Copilot platform, LLM-Tool Compiler achieves up to four times more parallel calls than existing methods, reducing token costs and latency by up to 40% and 12%, respectively.
Enhancing the Efficiency and Accuracy of Underlying Asset Reviews in Structured Finance: The Application of Multi-agent Framework
Wan, Xiangpeng, Deng, Haicheng, Zou, Kai, Xu, Shiqi
Structured finance, which involves restructuring diverse assets into securities like MBS, ABS, and CDOs, enhances capital market efficiency but presents significant due diligence challenges. This study explores the integration of artificial intelligence (AI) with traditional asset review processes to improve efficiency and accuracy in structured finance. Using both open-sourced and close-sourced large language models (LLMs), we demonstrate that AI can automate the verification of information between loan applications and bank statements effectively. While close-sourced models such as GPT-4 show superior performance, open-sourced models like LLAMA3 offer a cost-effective alternative. Dual-agent systems further increase accuracy, though this comes with higher operational costs. This research highlights AI's potential to minimize manual errors and streamline due diligence, suggesting a broader application of AI in financial document analysis and risk management.
Generative AI as a metacognitive agent: A comparative mixed-method study with human participants on ICF-mimicking exam performance
Pavlovic, Jelena, Krstic, Jugoslav, Mitrovic, Luka, Babic, Djordje, Milosavljevic, Adrijana, Nikolic, Milena, Karaklic, Tijana, Mitrovic, Tijana
Generative AI as a metacognitive agent: A comparative mixed-method study with human participants on ICF-mimicking exam performance Jelena Pavlović University of Belgrade, Faculty of Philosophy & Koučing centar Resarch Lab Jugoslav Krstić, Luka Mitrović, Đorđe Babić, Adrijana Milosavljević, Milena Nikolić, Tijana Karaklić & Tijana Mitrović Koučing centar Research Lab Abstract This study investigates the metacognitive capabilities of Large Language Models (LLMs) relative to human metacognition in the context of the International Coaching Federation (ICF)-mimicking exam, a situational judgment test related to coaching competencies. Using a mixed-method approach, we assessed the metacognitive performance--including sensitivity, accuracy in probabilistic predictions, and bias--of human participants and five advanced LLMs: GPT-4, Claude-3-Opus 3, Mistral Large, Llama 3, and Gemini 1.5 Pro. The results indicate that LLMs outperformed humans across all metacognitive metrics, particularly in terms of reduced overconfidence, compared to humans. However, both LLMs and humans showed less adaptability in ambiguous scenarios, adhering closely to predefined decision frameworks. The study suggests that Generative AI can effectively engage in human-like metacognitive processing without conscious awareness. Implications of the study are discussed in relation to development of AI simulators that scaffold cognitive and metacognitive aspects of mastering coaching competencies. More broadly, implications of these results are discussed in relation to development of metacognitive modules that lead towards more autonomous and intuitive AI systems. Keywords: Generative AI, metacognition, metacognitive agents, ICF exam Introduction Metacognition, the ability to understand and regulate one's cognitive processes, is a fundamental aspect of human learning, decision making and problem solving. Traditionally viewed as a conscious process, metacognition involves activities such as planning, monitoring, and evaluating one's performance during cognitive tasks. However, recent studies suggest that certain metacognitive processes can occur without conscious awareness, challenging the traditional boundaries of how metacognition is understood and measured Kentridge and Heywood (2000). In the field of generative artificial intelligence, particularly in Large Language Models (LLMs), metacognitive-like processes may manifest as algorithms adapt, learn, and optimize performance. This raises intriguing questions about the nature of metacognition in non-conscious entities and its comparison to human metacognitive processes. The present study aims to explore these questions by comparing the metacognitive processes of human participants and LLMs within the context of the International Coaching Federation (ICF) exam performance.
On the Foundations of Earth and Climate Foundation Models
Zhu, Xiao Xiang, Xiong, Zhitong, Wang, Yi, Stewart, Adam J., Heidler, Konrad, Wang, Yuanyuan, Yuan, Zhenghang, Dujardin, Thomas, Xu, Qingsong, Shi, Yilei
These authors contributed equally to this work. Abstract Foundation models have enormous potential in advancing Earth and climate sciences, however, current approaches may not be optimal as they focus on a few basic features of a desirable Earth and climate foundation model. Crafting the ideal Earth foundation model, we define eleven features which would allow such a foundation model to be beneficial for any geoscientific downstream application in an environmental-and human-centric manner. We further shed light on the way forward to achieve the ideal model and to evaluate Earth foundation models. What comes after foundation models? Energy efficient adaptation, adversarial defenses, and interpretability are among the emerging directions. In the past decade in particular, we have witnessed a paradigm shift from single-purpose models to general-purpose models, and from supervised pre-training to self-supervised pre-training. The majority of FMs like CLIP and GPT focus on the image and text domains. In this work, we specifically focus on "data" and "downstream tasks" relating to the Earth and its climate system, as shown in Figure 1. We choose to limit the scope of our work to the Earth's surface and atmosphere for three reasons. First, the Earth's surface and troposphere are our home, and include the majority of processes that directly impact and are impacted by human activity.
ESP: Extro-Spective Prediction for Long-term Behavior Reasoning in Emergency Scenarios
Wang, Dingrui, Lai, Zheyuan, Li, Yuda, Wu, Yi, Ma, Yuexin, Betz, Johannes, Yang, Ruigang, Li, Wei
Emergent-scene safety is the key milestone for fully autonomous driving, and reliable on-time prediction is essential to maintain safety in emergency scenarios. However, these emergency scenarios are long-tailed and hard to collect, which restricts the system from getting reliable predictions. In this paper, we build a new dataset, which aims at the long-term prediction with the inconspicuous state variation in history for the emergency event, named the Extro-Spective Prediction (ESP) problem. Based on the proposed dataset, a flexible feature encoder for ESP is introduced to various prediction methods as a seamless plug-in, and its consistent performance improvement underscores its efficacy. Furthermore, a new metric named clamped temporal error (CTE) is proposed to give a more comprehensive evaluation of prediction performance, especially in time-sensitive emergency events of subseconds. Interestingly, as our ESP features can be described in human-readable language naturally, the application of integrating into ChatGPT also shows huge potential. The ESP-dataset and all benchmarks are released at https://dingrui-wang.github.io/ESP-Dataset/.
Zero-shot LLM-guided Counterfactual Generation for Text
Bhattacharjee, Amrita, Moraffah, Raha, Garland, Joshua, Liu, Huan
Counterfactual examples are frequently used for model development and evaluation in many natural language processing (NLP) tasks. Although methods for automated counterfactual generation have been explored, such methods depend on models such as pre-trained language models that are then fine-tuned on auxiliary, often task-specific datasets. Collecting and annotating such datasets for counterfactual generation is labor intensive and therefore, infeasible in practice. Therefore, in this work, we focus on a novel problem setting: \textit{zero-shot counterfactual generation}. To this end, we propose a structured way to utilize large language models (LLMs) as general purpose counterfactual example generators. We hypothesize that the instruction-following and textual understanding capabilities of recent LLMs can be effectively leveraged for generating high quality counterfactuals in a zero-shot manner, without requiring any training or fine-tuning. Through comprehensive experiments on various downstream tasks in natural language processing (NLP), we demonstrate the efficacy of LLMs as zero-shot counterfactual generators in evaluating and explaining black-box NLP models.
GPT-Enabled Cybersecurity Training: A Tailored Approach for Effective Awareness
Al-Dhamari, Nabil, Clarke, Nathan
This study explores the limitations of traditional Cybersecurity Awareness and Training (CSAT) programs and proposes an innovative solution using Generative Pre-Trained Transformers (GPT) to address these shortcomings. Traditional approaches lack personalization and adaptability to individual learning styles. To overcome these challenges, the study integrates GPT models to deliver highly tailored and dynamic cybersecurity learning expe-riences. Leveraging natural language processing capabilities, the proposed approach personalizes training modules based on individual trainee pro-files, helping to ensure engagement and effectiveness. An experiment using a GPT model to provide a real-time and adaptive CSAT experience through generating customized training content. The findings have demonstrated a significant improvement over traditional programs, addressing issues of en-gagement, dynamicity, and relevance. GPT-powered CSAT programs offer a scalable and effective solution to enhance cybersecurity awareness, provid-ing personalized training content that better prepares individuals to miti-gate cybersecurity risks in their specific roles within the organization.
MEDVOC: Vocabulary Adaptation for Fine-tuning Pre-trained Language Models on Medical Text Summarization
Balde, Gunjan, Roy, Soumyadeep, Mondal, Mainack, Ganguly, Niloy
This work presents a dynamic vocabulary adaptation strategy, MEDVOC, for fine-tuning pre-trained language models (PLMs) like BertSumAbs, BART, and PEGASUS for improved medical text summarization. In contrast to existing domain adaptation approaches in summarization, MEDVOC treats vocabulary as an optimizable parameter and optimizes the PLM vocabulary based on fragment score conditioned only on the downstream task's reference summaries. Unlike previous works on vocabulary adaptation (limited only to classification tasks), optimizing vocabulary based on summarization tasks requires an extremely costly intermediate fine-tuning step on large summarization datasets. To that end, our novel fragment score-based hyperparameter search very significantly reduces this fine-tuning time -- from 450 days to less than 2 days on average. Furthermore, while previous works on vocabulary adaptation are often primarily tied to single PLMs, MEDVOC is designed to be deployable across multiple PLMs (with varying model vocabulary sizes, pre-training objectives, and model sizes) -- bridging the limited vocabulary overlap between the biomedical literature domain and PLMs. MEDVOC outperforms baselines by 15.74% in terms of Rouge-L in zero-shot setting and shows gains of 17.29% in high Out-Of-Vocabulary (OOV) concentrations. Our human evaluation shows MEDVOC generates more faithful medical summaries (88% compared to 59% in baselines). We make the codebase publicly available at https://github.com/gb-kgp/MEDVOC.
Evaluating Text Summaries Generated by Large Language Models Using OpenAI's GPT
Shakil, Hassan, Mahi, Atqiya Munawara, Nguyen, Phuoc, Ortiz, Zeydy, Mardini, Mamoun T.
In the contemporary era characterized by a deluge of data, the intelligence community faces the challenge of information overload, needing to process vast amounts of information swiftly and effectively. The ability to generate succinct, clear, and actionable summaries from diverse data sources is crucial, as it often determines the success of strategic objectives in this information-rich environment. As the demand for systems capable of automating large-scale text summarization without compromising on quality or relevance intensifies, the role of such technologies becomes increasingly critical Liu and Lapata [2019]. Text summarization, a pivotal task within Natural Language Processing (NLP), has found widespread application across various domains, including news aggregation and the distillation of extensive documents into manageable summaries. The exponential growth in data underscores the utility of text summarization in enhancing content accessibility and comprehension, thus facilitating more efficient navigation through information landscapes Chouikhi and Alsuhaibani [2022].