Africa
Calibrating Large Language Models Using Their Generations Only
Ulmer, Dennis, Gubri, Martin, Lee, Hwaran, Yun, Sangdoo, Oh, Seong Joon
As large language models (LLMs) are increasingly deployed in user-facing applications, building trust and maintaining safety by accurately quantifying a model's confidence in its prediction becomes even more important. However, finding effective ways to calibrate LLMs - especially when the only interface to the models is their generated text - remains a challenge. We propose APRICOT (auxiliary prediction of confidence targets): A method to set confidence targets and train an additional model that predicts an LLM's confidence based on its textual input and output alone. This approach has several advantages: It is conceptually simple, does not require access to the target model beyond its output, does not interfere with the language generation, and has a multitude of potential usages, for instance by verbalizing the predicted confidence or adjusting the given answer based on the confidence. We show how our approach performs competitively in terms of calibration error for white-box and black-box LLMs on closed-book question-answering to detect incorrect LLM answers.
MaiBaam Annotation Guidelines
Blaschke, Verena, Kovaฤiฤ, Barbara, Peng, Siyao, Plank, Barbara
This document provides annotation guidelines for MaiBaam, a Bavarian corpus annotated with part-of-speech (POS) tags and syntactic dependencies. MaiBaam belongs to the Universal Dependencies (UD) project (Zeman et al., 2023; de Marneffe et al., 2021), and our annotations elaborate on the general and German UD version 2 guidelines. This document is structured broadly in the order we prepare and annotate sentences: first, preprocessing and tokenization ( 1), then general recaps of POS tags ( 2) and dependencies ( 3), before we go into annotation decisions that would also apply to German ( 4) and lastly decisions that are specific to Bavarian grammar ( 5). Many examples are written in German, since the standardized orthography makes it easier to search this PDF. We only annotate UD-style POS tags (UPOS tags) and dependencies and add the SpaceAfter=No feature where appropriate, but do not add any other information (no lemma, XPOS tags, morphological features, enhanced dependencies or miscellaneous annotations). This document is primarily directed at present and future annotators of MaiBaam. We publish it to additionally allow others working with MaiBaam or annotating similar data to better understand the decisions we have made.
Diffusion Lens: Interpreting Text Encoders in Text-to-Image Pipelines
Toker, Michael, Orgad, Hadas, Ventura, Mor, Arad, Dana, Belinkov, Yonatan
Text-to-image diffusion models (T2I) use a latent representation of a text prompt to guide the image generation process. However, the process by which the encoder produces the text representation is unknown. We propose the Diffusion Lens, a method for analyzing the text encoder of T2I models by generating images from its intermediate representations. Using the Diffusion Lens, we perform an extensive analysis of two recent T2I models. Exploring compound prompts, we find that complex scenes describing multiple objects are composed progressively and more slowly compared to simple scenes; Exploring knowledge retrieval, we find that representation of uncommon concepts requires further computation compared to common concepts, and that knowledge retrieval is gradual across layers. Overall, our findings provide valuable insights into the text encoder component in T2I pipelines.
End-to-end solution for linked open data query logs analytics
In big data Era, significant advances in e-commerce, targeted marketing, social shopping, e-tourism, etc. are derived basically from collective intelligence. Such applications mainly exploit data generated by users to extract different valuable information. User content represents data, information, or media content voluntarily provided by people Krumm et al. [2008], when they interact with web sites, social media, and data sources, etc. This data regroups social data, YouTube videos, blogs and micro-blogs, query-logs, etc. Analysis of this data provides useful information helping to understand user behavior, user opinions, topics of interest, etc. It helps to detect hidden patterns and to construct users' profiles, in order to propose user-centric solutions like: recommendation systems, content personalization, cache improvement, etc. for successful user experience.
Stephen Salter obituary
Stephen Salter, who has died aged 85, was the inventor of the Salter's Duck, a wave-power device that was the first of its kind and promised to provide a new source of renewable energy for the world โ until it was effectively killed off by the nuclear industry. In 1982, after eight years of development under Salter's direction at Edinburgh University, the United Kingdom Atomic Energy Authority (UKAEA) was asked by the government to see if the duck might be a cost-effective way of making large quantities of electricity. To the great surprise of Salter, and others, the UKAEA came to the conclusion that it was uneconomic, and that no further government funding should be given to the project. A decade later it emerged that thanks to a misplaced decimal point, the review had made Salter's duck look 10 times more expensive than the experiments showed it was likely to be. The UKAEA claimed this was just a mistake, but Salter, who had never been allowed to see the results of the secret evaluation, put it another way: asking the nuclear industry to evaluate an alternative source of energy was like putting King Herod in charge of a children's home, he suggested.
Happy International Women's Day!
To celebrate International Women's Day, we take a look back over the past 12 months and highlight some of the women we've interviewed and featured, and who've written about their research on AIhub. Elizabeth Ondula is an Electrical Engineer from the Technical University of Kenya and is currently a PhD student of Computer Science at USC. She is a member of the Autonomous Networks Research Group, and co-organizes a bi-weekly reinforcement learning group, SUITERS-RL. Prior to academia, she had roles as a Software Engineer at IBM Research in Kenya, Head of Product Development of Brave Venture Labs and Co-lead of Hardware Research at iHub Nairobi. We interviewed Elizabeth as part of our series featuring the AAAI Doctoral Consortium participants.
Credit Card Fraud Detection in the Nigerian Financial Sector: A Comparison of Unsupervised TensorFlow-Based Anomaly Detection Techniques, Autoencoders and PCA Algorithm
Credit card fraud is a major cause of national concern in the Nigerian financial sector, affecting hundreds of transactions per second and impacting international e-commerce negatively. Despite the rapid spread and adoption of online marketing, millions of Nigerians are prevented from transacting in several countries with local credit cards due to bans and policies directed at restricting credit card fraud. Presently, a myriad of technologies exist to detect fraudulent transactions, a few of which are adopted by Nigerian financial institutions to proactively manage the situation. Fraud detection allows institutions to restrict offenders from networks and with a centralized banking identity management system, such as the Bank Verification Number used by the Central Bank of Nigeria, offenders who may have stolen other people's identities can be back-traced and their bank accounts frozen. This paper aims to compare the effectiveness of two fraud detection technologies that are projected to work fully independent of human intervention to possibly predict and detect fraudulent credit card transactions. Autoencoders as an Unsupervised Tensorflow-Based Anomaly Detection Technique generally offers greater performance in dimensionality reduction than the Principal Component Analysis, and this theory was tested out on Nigerian credit card transaction data. Results demonstrate that autoencoders are better suited to analyzing complex and extensive datasets and offer more reliable results with minimal mislabeling than the PCA algorithm.
SeeGULL Multilingual: a Dataset of Geo-Culturally Situated Stereotypes
Bhutani, Mukul, Robinson, Kevin, Prabhakaran, Vinodkumar, Dave, Shachi, Dev, Sunipa
While generative multilingual models are rapidly being deployed, their safety and fairness evaluations are largely limited to resources collected in English. This is especially problematic for evaluations targeting inherently socio-cultural phenomena such as stereotyping, where it is important to build multi-lingual resources that reflect the stereotypes prevalent in respective language communities. However, gathering these resources, at scale, in varied languages and regions pose a significant challenge as it requires broad socio-cultural knowledge and can also be prohibitively expensive. To overcome this critical gap, we employ a recently introduced approach that couples LLM generations for scale with culturally situated validations for reliability, and build SeeGULL Multilingual, a global-scale multilingual dataset of social stereotypes, containing over 25K stereotypes, spanning 20 languages, with human annotations across 23 regions, and demonstrate its utility in identifying gaps in model evaluations. Content warning: Stereotypes shared in this paper can be offensive.
A Survey on Knowledge Distillation of Large Language Models
Xu, Xiaohan, Li, Ming, Tao, Chongyang, Shen, Tao, Cheng, Reynold, Li, Jinyang, Xu, Can, Tao, Dacheng, Zhou, Tianyi
In the era of Large Language Models (LLMs), Knowledge Distillation (KD) emerges as a pivotal methodology for transferring advanced capabilities from leading proprietary LLMs, such as GPT-4, to their open-source counterparts like LLaMA and Mistral. Additionally, as open-source LLMs flourish, KD plays a crucial role in both compressing these models, and facilitating their self-improvement by employing themselves as teachers. This paper presents a comprehensive survey of KD's role within the realm of LLM, highlighting its critical function in imparting advanced knowledge to smaller models and its utility in model compression and self-improvement. Our survey is meticulously structured around three foundational pillars: \textit{algorithm}, \textit{skill}, and \textit{verticalization} -- providing a comprehensive examination of KD mechanisms, the enhancement of specific cognitive abilities, and their practical implications across diverse fields. Crucially, the survey navigates the intricate interplay between data augmentation (DA) and KD, illustrating how DA emerges as a powerful paradigm within the KD framework to bolster LLMs' performance. By leveraging DA to generate context-rich, skill-specific training data, KD transcends traditional boundaries, enabling open-source models to approximate the contextual adeptness, ethical alignment, and deep semantic insights characteristic of their proprietary counterparts. This work aims to provide an insightful guide for researchers and practitioners, offering a detailed overview of current methodologies in KD and proposing future research directions. Importantly, we firmly advocate for compliance with the legal terms that regulate the use of LLMs, ensuring ethical and lawful application of KD of LLMs. An associated Github repository is available at https://github.com/Tebmer/Awesome-Knowledge-Distillation-of-LLMs.
Tapilot-Crossing: Benchmarking and Evolving LLMs Towards Interactive Data Analysis Agents
Li, Jinyang, Huo, Nan, Gao, Yan, Shi, Jiayi, Zhao, Yingxiu, Qu, Ge, Wu, Yurong, Ma, Chenhao, Lou, Jian-Guang, Cheng, Reynold
Interactive Data Analysis, the collaboration between humans and LLM agents, enables real-time data exploration for informed decision-making. The challenges and costs of collecting realistic interactive logs for data analysis hinder the quantitative evaluation of Large Language Model (LLM) agents in this task. To mitigate this issue, we introduce Tapilot-Crossing, a new benchmark to evaluate LLM agents on interactive data analysis. Tapilot-Crossing contains 1024 interactions, covering 4 practical scenarios: Normal, Action, Private, and Private Action. Notably, Tapilot-Crossing is constructed by an economical multi-agent environment, Decision Company, with few human efforts. We evaluate popular and advanced LLM agents in Tapilot-Crossing, which underscores the challenges of interactive data analysis. Furthermore, we propose Adaptive Interaction Reflection (AIR), a self-generated reflection strategy that guides LLM agents to learn from successful history. Experiments demonstrate that Air can evolve LLMs into effective interactive data analysis agents, achieving a relative performance improvement of up to 44.5%.