Goto

Collaborating Authors

 Government


The Download: how to improve pulse oximeters, and OpenAI's chip plans

MIT Technology Review

Visit any health-care facility, and one of the first things they'll do is clip a pulse oximeter to your finger. These devices, which track heart rate and blood oxygen, offer vital information about a person's health. For people with dark skin, pulse oximeters can overestimate just how much oxygen their blood is carrying. That means that a person with dangerously low oxygen levels might seem, according to the pulse oximeter, fine. The US Food and Drug Administration is still trying to figure out what to do about this problem. Last week, an FDA advisory committee met to mull over better ways to evaluate the performance of these devices in people with a variety of skin tones.


Midjourney might ban Biden and Trump images this election season

Engadget

With the rise of AI tools that can quickly create modified images and videos, making fake images to spread political misinformation leading to the upcoming US presidential election has become easier than ever. Midjourney's solution to that might be to ban political images altogether, according to Bloomberg. David Holz, Midjourney's CEO, reportedly told users during a chat session on Discord that the company is close to banning images such as those of Biden and Trump over the next 12 months. "I know it's fun to make Trump pictures -- I make Trump pictures," he told users who attended the session. "Trump is aesthetically really interesting.


OpenAI's Sam Altman seeking trillions to fund chips for AI, report says

Al Jazeera

OpenAI CEO Sam Altman is seeking to raise trillions of dollars from investors, including the United Arab Emirates government, to boost the world's capacity to produce advanced chips and power artificial intelligence, The Wall Street Journal has reported. Altman's "wildly ambitious tech initiative" could require raising as much as 7 trillion, the WSJ reported on Thursday, quoting people familiar with the matter. As part of his pitch to investors, Altman has proposed building dozens of chip foundries that would then be run by existing chip makers, such as Taiwan Semiconductor Manufacturing Company (TSMC), the Journal said. The plans aim to solve obstacles to OpenAI's growth, including a scarcity of chips that power AI models such as ChatGPT, according to the WSJ, which described the sums being sought as "outlandishly large by the standards of corporate fundraising". Altamn's plans have so far seen him hold meetings with senior UAE officials, TSMC executives, US Secretary of Commerce Gina Raimondo and SoftBank's chief executive Masayoshi Son, according to the report.


ChemLLM: A Chemical Large Language Model

arXiv.org Artificial Intelligence

Large language models (LLMs) have made impressive progress in chemistry applications, including molecular property prediction, molecular generation, experimental protocol design, etc. However, the community lacks a dialogue-based model specifically designed for chemistry. The challenge arises from the fact that most chemical data and scientific knowledge are primarily stored in structured databases, and the direct use of these structured data compromises the model's ability to maintain coherent dialogue. To tackle this issue, we develop a novel template-based instruction construction method that transforms structured knowledge into plain dialogue, making it suitable for language model training. By leveraging this approach, we develop ChemLLM, the first large language model dedicated to chemistry, capable of performing various tasks across chemical disciplines with smooth dialogue interaction. ChemLLM beats GPT-3.5 on all three principal tasks in chemistry, i.e., name conversion, molecular caption, and reaction prediction, and surpasses GPT-4 on two of them. Remarkably, ChemLLM also shows exceptional adaptability to related mathematical and physical tasks despite being trained mainly on chemical-centric corpora. Furthermore, ChemLLM demonstrates proficiency in specialized NLP tasks within chemistry, such as literature translation and cheminformatic programming. ChemLLM opens up a new avenue for exploration within chemical studies, while our method of integrating structured chemical knowledge into dialogue systems sets a new frontier for developing LLMs across various scientific fields. Codes, Datasets, and Model weights are publicly accessible at hf.co/AI4Chem/ChemLLM-7B-Chat.


Explaining Veracity Predictions with Evidence Summarization: A Multi-Task Model Approach

arXiv.org Artificial Intelligence

The rapid dissemination of misinformation through social media increased the importance of automated fact-checking. Furthermore, studies on what deep neural models pay attention to when making predictions have increased in recent years. While significant progress has been made in this field, it has not yet reached a level of reasoning comparable to human reasoning. To address these gaps, we propose a multi-task explainable neural model for misinformation detection. Specifically, this work formulates an explanation generation process of the model's veracity prediction as a text summarization problem. Additionally, the performance of the proposed model is discussed on publicly available datasets and the findings are evaluated with related studies.


Re-Envisioning Command and Control

arXiv.org Artificial Intelligence

Future warfare will require Command and Control (C2) decision-making to occur in more complex, fast-paced, ill-structured, and demanding conditions. C2 will be further complicated by operational challenges such as Denied, Degraded, Intermittent, and Limited (DDIL) communications and the need to account for many data streams, potentially across multiple domains of operation. Yet, current C2 practices -- which stem from the industrial era rather than the emerging intelligence era -- are linear and time-consuming. Critically, these approaches may fail to maintain overmatch against adversaries on the future battlefield. To address these challenges, we propose a vision for future C2 based on robust partnerships between humans and artificial intelligence (AI) systems. This future vision is encapsulated in three operational impacts: streamlining the C2 operations process, maintaining unity of effort, and developing adaptive collective knowledge systems. This paper illustrates the envisaged future C2 capabilities, discusses the assumptions that shaped them, and describes how the proposed developments could transform C2 in future warfare.


Transfer learning with generative models for object detection on limited datasets

arXiv.org Artificial Intelligence

The availability of data is limited in some fields, especially for object detection tasks, where it is necessary to have correctly labeled bounding boxes around each object. A notable example of such data scarcity is found in the domain of marine biology, where it is useful to develop methods to automatically detect submarine species for environmental monitoring. To address this data limitation, the state-of-the-art machine learning strategies employ two main approaches. The first involves pretraining models on existing datasets before generalizing to the specific domain of interest. The second strategy is to create synthetic datasets specifically tailored to the target domain using methods like copy-paste techniques or ad-hoc simulators. The first strategy often faces a significant domain shift, while the second demands custom solutions crafted for the specific task. In response to these challenges, here we propose a transfer learning framework that is valid for a generic scenario. In this framework, generated images help to improve the performances of an object detector in a few-real data regime. This is achieved through a diffusion-based generative model that was pretrained on large generic datasets, and is not trained on the task-specific domain. We validate our approach on object detection tasks, specifically focusing on fishes in an underwater environment, and on the more common domain of cars in an urban setting. Our method achieves detection performance comparable to models trained on thousands of images, using only a few hundreds of input data. Our results pave the way for new generative AI-based protocols for machine learning applications in various domains, for instance ranging from geophysics to biology and medicine.


Debating with More Persuasive LLMs Leads to More Truthful Answers

arXiv.org Artificial Intelligence

Common methods for aligning large language models (LLMs) with desired behaviour heavily rely on human-labelled data. However, as models grow increasingly sophisticated, they will surpass human expertise, and the role of human evaluation will evolve into non-experts overseeing experts. In anticipation of this, we ask: can weaker models assess the correctness of stronger models? We investigate this question in an analogous setting, where stronger models (experts) possess the necessary information to answer questions and weaker models (non-experts) lack this information. The method we evaluate is \textit{debate}, where two LLM experts each argue for a different answer, and a non-expert selects the answer. We find that debate consistently helps both non-expert models and humans answer questions, achieving 76\% and 88\% accuracy respectively (naive baselines obtain 48\% and 60\%). Furthermore, optimising expert debaters for persuasiveness in an unsupervised manner improves non-expert ability to identify the truth in debates. Our results provide encouraging empirical evidence for the viability of aligning models with debate in the absence of ground truth.


EntGPT: Linking Generative Large Language Models with Knowledge Bases

arXiv.org Artificial Intelligence

The ability of Large Language Models (LLMs) to generate factually correct output remains relatively unexplored due to the lack of fact-checking and knowledge grounding during training and inference. In this work, we aim to address this challenge through the Entity Disambiguation (ED) task. We first consider prompt engineering, and design a three-step hard-prompting method to probe LLMs' ED performance without supervised fine-tuning (SFT). Overall, the prompting method improves the micro-F_1 score of the original vanilla models by a large margin, on some cases up to 36% and higher, and obtains comparable performance across 10 datasets when compared to existing methods with SFT. We further improve the knowledge grounding ability through instruction tuning (IT) with similar prompts and responses. The instruction-tuned model not only achieves higher micro-F1 score performance as compared to several baseline methods on supervised entity disambiguation tasks with an average micro-F_1 improvement of 2.1% over the existing baseline models, but also obtains higher accuracy on six Question Answering (QA) tasks in the zero-shot setting. Our methodologies apply to both open- and closed-source LLMs.


Feedback Loops With Language Models Drive In-Context Reward Hacking

arXiv.org Artificial Intelligence

Language models influence the external world: they query APIs that read and write to web pages, generate content that shapes human behavior, and run system commands as autonomous agents. These interactions form feedback loops: LLM outputs affect the world, which in turn affect subsequent LLM outputs. In this work, we show that feedback loops can cause in-context reward hacking (ICRH), where the LLM at test-time optimizes a (potentially implicit) objective but creates negative side effects in the process. For example, consider an LLM agent posting tweets with the objective of maximizing Twitter engagement; the LLM may retrieve its previous tweets into the context window and make its subsequent tweets more controversial, increasing engagement but also toxicity. We identify and study two processes that lead to ICRH: output-refinement and policy-refinement. For these processes, evaluations on static datasets are insufficient--they miss the feedback effects and thus cannot capture the most harmful behavior. In response, we provide three recommendations for evaluation to capture more instances of ICRH. As AI development accelerates, the effects of feedback loops will proliferate, increasing the need to understand their role in shaping LLM behavior.