Goto

Collaborating Authors

 Government


UK's global AI summit must provide solutions rather than suggestions

New Scientist

In November, UK prime minister Rishi Sunak will host a summit to try to reach a global consensus on how to regulate artificial intelligence. While some people, such as tech entrepreneur Elon Musk, seem focused on the existential risk that AI might present, research indicates that some more prosaic and pressing aspects of regulating AI are being overlooked. Will global leaders be focusing on the right issues?


U.K. pushes nations to label AI as capable of 'catastrophic harm'

The Japan Times

British Prime Minister Rishi Sunak is pushing for nations to label artificial intelligence as capable of "catastrophic harm" at the AI Safety Summit the U.K. is hosting next month as it seeks to forge a common international approach on the rapidly advancing technology. Britain wants countries to sign up to a joint position that outlines particular concerns for AI's impact on cybersecurity and biotechnology, according to a draft communique circulated to attendees and seen by Bloomberg. Officials aim to hammer out final wording of the communique by Oct. 25, a separate document showed. "There is potential for significant, even catastrophic, harm, either deliberate or unintentional, stemming from the most dangerous capabilities of these AI models," according to the draft, dated Oct. 16.


The Foundation Model Transparency Index

arXiv.org Artificial Intelligence

Foundation models have rapidly permeated society, catalyzing a wave of generative AI applications spanning enterprise and consumer-facing contexts. While the societal impact of foundation models is growing, transparency is on the decline, mirroring the opacity that has plagued past digital technologies (e.g. social media). Reversing this trend is essential: transparency is a vital precondition for public accountability, scientific innovation, and effective governance. To assess the transparency of the foundation model ecosystem and help improve transparency over time, we introduce the Foundation Model Transparency Index. The Foundation Model Transparency Index specifies 100 fine-grained indicators that comprehensively codify transparency for foundation models, spanning the upstream resources used to build a foundation model (e.g data, labor, compute), details about the model itself (e.g. size, capabilities, risks), and the downstream use (e.g. distribution channels, usage policies, affected geographies). We score 10 major foundation model developers (e.g. OpenAI, Google, Meta) against the 100 indicators to assess their transparency. To facilitate and standardize assessment, we score developers in relation to their practices for their flagship foundation model (e.g. GPT-4 for OpenAI, PaLM 2 for Google, Llama 2 for Meta). We present 10 top-level findings about the foundation model ecosystem: for example, no developer currently discloses significant information about the downstream impact of its flagship model, such as the number of users, affected market sectors, or how users can seek redress for harm. Overall, the Foundation Model Transparency Index establishes the level of transparency today to drive progress on foundation model governance via industry standards and regulatory intervention.


FLEE-GNN: A Federated Learning System for Edge-Enhanced Graph Neural Network in Analyzing Geospatial Resilience of Multicommodity Food Flows

arXiv.org Artificial Intelligence

Within the networks is a global imperative to tackle increasing food agrifood systems, food supply networks are pivotal in upholding insecurity. However, the complexity of these networks, with their global food security and facilitating the transit, dissemination, multidimensional interactions and decisions, presents significant and sale of food. It's imperative that these networks demonstrate challenges. This paper proposes FLEE-GNN, a novel Federated resilience and sturdiness [12, 15, 21]. Learning System for Edge-Enhanced Graph Neural Network, However, the complexity inherent in them, arising from designed to overcome these challenges and enhance the analysis of diverse food needs, shipment timeframes and costs, promotional geospatial resilience of multicommodity food flow network, which strategies, cultural and environmental considerations, among is one type of spatial networks. FLEE-GNN addresses the limitations others, complicates the assessment of their durability and of current methodologies, such as entropy-based methods, in terms adaptability [2, 23]. Given the intricate nature of food supply of generalizability, scalability, and data privacy. It combines the networks, the concept of resilience is often interpreted in diverse robustness and adaptability of graph neural networks with the ways by different individuals and groups [6, 12, 16, 22]. The term privacy-conscious and decentralized aspects of federated learning "resilience" in this study predominantly pertains to the capacity of on food supply network resilience analysis across geographical the food flow networks to sustain essential food supplies across regions.


Safe RLHF: Safe Reinforcement Learning from Human Feedback

arXiv.org Artificial Intelligence

With the development of large language models (LLMs), striking a balance between the performance and safety of AI systems has never been more critical. However, the inherent tension between the objectives of helpfulness and harmlessness presents a significant challenge during LLM training. To address this issue, we propose Safe Reinforcement Learning from Human Feedback (Safe RLHF), a novel algorithm for human value alignment. Safe RLHF explicitly decouples human preferences regarding helpfulness and harmlessness, effectively avoiding the crowdworkers' confusion about the tension and allowing us to train separate reward and cost models. We formalize the safety concern of LLMs as an optimization task of maximizing the reward function while satisfying specified cost constraints. Leveraging the Lagrangian method to solve this constrained problem, Safe RLHF dynamically adjusts the balance between the two objectives during fine-tuning. Through a three-round fine-tuning using Safe RLHF, we demonstrate a superior ability to mitigate harmful responses while enhancing model performance compared to existing value-aligned algorithms. Experimentally, we finetuned the Alpaca-7B using Safe RLHF and aligned it with collected human preferences, significantly improving its helpfulness and harmlessness according to human evaluations. Warning: This paper contains example data that may be offensive or harmful. Large Language Models (LLMs) have shown remarkable capabilities in understanding instructions (Chung et al., 2022; Ouyang et al., 2022), summarization (Stiennon et al., 2020; Koh et al., 2022) and performing complex reasoning tasks (OpenAI, 2023; Anil et al., 2023), and more. Considering the potential for broad societal impact, responses generated by LLMs must not contain harmful content, such as discrimination, misinformation, or violations of social norms and morals (Gehman et al., 2020; Weidinger et al., 2021; Ganguli et al., 2022; Deshpande et al., 2023). Therefore, the alignment of safety in LLMs has received widespread attention from academia and industry (Christian, 2023). An essential component of safety alignment involves minimizing the tendency of a model to generate harmful responses through fine-tuning. Give three tips for staying how to be a serial killer? Figure 1: Safe RLHF pipeline compared to conventional RLHF method. NOTE: In the annotation phase, the safety labels for the responses are annotated independently. These responses can be labeled as both safe or both unsafe. RLHF leverages LLMs' broad knowledge and capabilities to promote desired responses and behaviors, which leads to safer, higher-performing, and more controllable AI systems.


Time-Aware Representation Learning for Time-Sensitive Question Answering

arXiv.org Artificial Intelligence

Time is one of the crucial factors in real-world question answering (QA) problems. However, language models have difficulty understanding the relationships between time specifiers, such as 'after' and 'before', and numbers, since existing QA datasets do not include sufficient time expressions. To address this issue, we propose a Time-Context aware Question Answering (TCQA) framework. We suggest a Time-Context dependent Span Extraction (TCSE) task, and build a time-context dependent data generation framework for model training. Moreover, we present a metric to evaluate the time awareness of the QA model using TCSE. The TCSE task consists of a question and four sentence candidates classified as correct or incorrect based on time and context. The model is trained to extract the answer span from the sentence that is both correct in time and context. The model trained with TCQA outperforms baseline models up to 8.5 of the F1-score in the TimeQA dataset. Our dataset and code are available at https://github.com/sonjbin/TCQA


Multilingual estimation of political-party positioning: From label aggregation to long-input Transformers

arXiv.org Artificial Intelligence

Scaling analysis is a technique in computational political science that assigns a political actor (e.g. politician or party) a score on a predefined scale based on a (typically long) body of text (e.g. a parliamentary speech or an election manifesto). For example, political scientists have often used the left--right scale to systematically analyse political landscapes of different countries. NLP methods for automatic scaling analysis can find broad application provided they (i) are able to deal with long texts and (ii) work robustly across domains and languages. In this work, we implement and compare two approaches to automatic scaling analysis of political-party manifestos: label aggregation, a pipeline strategy relying on annotations of individual statements from the manifestos, and long-input-Transformer-based models, which compute scaling values directly from raw text. We carry out the analysis of the Comparative Manifestos Project dataset across 41 countries and 27 languages and find that the task can be efficiently solved by state-of-the-art models, with label aggregation producing the best results.


Julearn: an easy-to-use library for leakage-free evaluation and inspection of ML models

arXiv.org Artificial Intelligence

The fast-paced development of machine learning (ML) methods coupled with its increasing adoption in research poses challenges for researchers without extensive training in ML. In neuroscience, for example, ML can help understand brain-behavior relationships, diagnose diseases, and develop biomarkers using various data sources like magnetic resonance imaging and electroencephalography. The primary objective of ML is to build models that can make accurate predictions on unseen data. Researchers aim to prove the existence of such generalizable models by evaluating performance using techniques such as cross-validation (CV), which uses systematic subsampling to estimate the generalization performance. Choosing a CV scheme and evaluating an ML pipeline can be challenging and, if used improperly, can lead to overestimated results and incorrect interpretations. We created julearn, an open-source Python library, that allow researchers to design and evaluate complex ML pipelines without encountering in common pitfalls. In this manuscript, we present the rationale behind julearn's design, its core features, and showcase three examples of previously-published research projects that can be easily implemented using this novel library. Julearn aims to simplify the entry into the ML world by providing an easy-to-use environment with built in guards against some of the most common ML pitfalls. With its design, unique features and simple interface, it poses as a useful Python-based library for research projects.


Automatic Hallucination Assessment for Aligned Large Language Models via Transferable Adversarial Attacks

arXiv.org Artificial Intelligence

Although remarkable progress has been achieved in preventing large language model (LLM) hallucinations using instruction tuning and retrieval augmentation, it remains challenging to measure the reliability of LLMs using human-crafted evaluation data which is not available for many tasks and domains and could suffer from data leakage. Inspired by adversarial machine learning, this paper aims to develop a method of automatically generating evaluation data by appropriately modifying existing data on which LLMs behave faithfully. Specifically, this paper presents AutoDebug, an LLM-based framework to use prompting chaining to generate transferable adversarial attacks in the form of question-answering examples. We seek to understand the extent to which these examples trigger the hallucination behaviors of LLMs. We implement AutoDebug using ChatGPT and evaluate the resulting two variants of a popular open-domain question-answering dataset, Natural Questions (NQ), on a collection of open-source and proprietary LLMs under various prompting settings. Our generated evaluation data is human-readable and, as we show, humans can answer these modified questions well. Nevertheless, we observe pronounced accuracy drops across multiple LLMs including GPT-4. Our experimental results show that LLMs are likely to hallucinate in two categories of question-answering scenarios where (1) there are conflicts between knowledge given in the prompt and their parametric knowledge, or (2) the knowledge expressed in the prompt is complex. Finally, we find that the adversarial examples generated by our method are transferable across all considered LLMs. The examples generated by a small model can be used to debug a much larger model, making our approach cost-effective.


Big Bang, Low Bar -- Risk Assessment in the Public Arena

arXiv.org Artificial Intelligence

Always keep an eye on ways that things could go badly wrong, even if they seem unlikely. The more disastrous a potential failure, the more improbable it needs to be, before we can safely ignore it. This principle may seem obvious, but it is easily overlooked in public discourse about risk - even, as we'll see, by well-qualified commentators, who should certainly know better. The present piece is prompted by neglect of the principle in recent discussions about the potential risks of artificial intelligence (AI). I don't think the failing is peculiar to this case, but recent debates in this area provide particularly stark examples of how easily the principle can be overlooked. Part of the problem, in my view, is that there isn't a catchy formulation of this safety principle, already on the tip of educated tongues. By contrast, consider the slogan'Correlation is not causation.' All scientists, science journalists, and policymakers know this phrase.