Large Language Model
OpenAI putting 'shiny products' above safety, says departing researcher
A former senior employee at OpenAI has said the company behind ChatGPT is prioritising "shiny products" over safety, revealing that he quit after a disagreement over key aims reached "breaking point". Jan Leike was a key safety researcher at OpenAI as its co-head of superalignment, ensuring that powerful artificial intelligence systems adhere to human values and aims. His intervention comes before a global artificial intelligence summit in Seoul next week, where politicians, experts and tech executives will discuss oversight of the technology. Leike resigned days after the San Francisco-based company launched its latest AI model, GPT-4o. His departure means two senior safety figures at OpenAI have left this week following the resignation of Ilya Sutskever, OpenAI's co-founder and fellow co-head of superalignment.
As the AI world gathers in Seoul, can an accelerating industry balance progress against safety?
This week, artificial intelligence caught up with the future โ or at least Hollywood's idea of it from a decade ago. "It feels like AI from the movies," wrote the OpenAI chief executive, Sam Altman, of his latest system, an impressive virtual assistant. To underline his point he posted a single word on X โ "her" โ referring to the 2013 film starring Joaquin Phoenix as a man who falls in love with a futuristic version of Siri or Alexa, voiced by Scarlett Johansson. For some experts, that new AI, GPT-4o, will be an unsettling reminder of their concerns about the technology's rapid advances, with a key OpenAI safety researcher leaving this week following a disagreement over the company's direction. For others the GPT-4o release will be confirmation that innovation continues in a field promising benefits for all. Next week's global AI summit in Seoul, attended by ministers, experts and tech executives, will hear both perspectives, as underlined by a safety report released before the meeting that referred to potential positives as well as numerous risks.
How to watch the Microsoft Build 2024 keynote live on May 21
Springtime means it's keynote season in the tech world, and in 2024, that means "time to show off your AI bona fides." Google and OpenAI have already revealed big new upgrades to Gemini and ChatGPT this month, and now it's time for Microsoft Build. The tech giant's annual developer conference kicks off with a keynote slated for Tuesday, May 21 at 12 PM ET/9 AM PT, and you can watch the entire event live on YouTube (which is also embedded below) and at Microsoft's site (registration required). What about that Microsoft Surface event you may have heard about? Don't worry, here's the tl;dr version of what to expect, summarized from our more in-depth What to expect from Microsoft Build 2024: The Surface event, Windows 11 and AI.
Identifying and Aligning Medical Claims Made on Social Media with Medical Evidence
Evidence-based medicine is the practise of making medical decisions that adhere to the latest, and best known evidence at that time. Currently, the best evidence is often found in the form of documents, such as randomized control trials, meta-analyses and systematic reviews. This research focuses on aligning medical claims made on social media platforms with this medical evidence. By doing so, individuals without medical expertise can more effectively assess the veracity of such medical claims. We study three core tasks: identifying medical claims, extracting medical vocabulary from these claims, and retrieving evidence relevant to those identified medical claims. We propose a novel system that can generate synthetic medical claims to aid each of these core tasks. We additionally introduce a novel dataset produced by our synthetic generator that, when applied to these tasks, demonstrates not only a more flexible and holistic approach, but also an improvement in all comparable metrics. We make our dataset, the Expansive Medical Claim Corpus (EMCC), available at https://zenodo.org/records/8321460. Keywords: Evidenced-based Medicine, PICO, Synthetic Generators, Information Retrieval
EnterpriseEM: Fine-tuned Embeddings for Enterprise Semantic Search
Rathinasamy, Kamalkumar, Nettar, Jayarama, Kumar, Amit, Manchanda, Vishal, Vijayakumar, Arun, Kataria, Ayush, Manjunath, Venkateshprasanna, GS, Chidambaram, Sodhi, Jaskirat Singh, Shaikh, Shoeb, Khan, Wasim Akhtar, Singh, Prashant, Ige, Tanishq Dattatray, Tiwari, Vipin, Mondal, Rajab Ali, K, Harshini, Reka, S, Amancharla, Chetana, Rahman, Faiz ur, A, Harikrishnan P, Saha, Indraneel, Tiwary, Bhavya, Patel, Navin Shankar, S, Pradeep T, J, Balaji A, Priyapravas, null, Tarafdar, Mohammed Rafee
In the context of enterprises accumulating proprietary unstructured data, AI-driven information retrieval solutions have emerged as vital tools for extracting relevant answers to employee queries. Traditional methods for developing such solutions often involve choosing between Retrieval Augmented Generation (RAG) or fine-tuned Large Language Models (LLMs). However, fine-tuned LLMs, comprising only generative models, lack a guarantee of factual accuracy, while RAG, comprising an embedding model and a generative model, assures factual precision (Lewis at al., 2020 [1]). Despite their superior performance in general, RAG based solutions often rely on pre-trained models, potentially leading to suboptimal alignment with enterprise-specific data. Addressing this challenge entails exploring two potential avenues: Firstly, recent studies such as RAFT (Zhang et al., 2024 [2]) explore the integration of fine-tuned generative models within a RAG pipeline to enhance accuracy, albeit requiring substantial domain-specific data to fine-tune the generative models. Alternatively, leveraging domain-specific embedding models within a RAG pipeline to enhance accuracy remains an underexplored area. Earlier efforts, such as BioBERT (Lee et al., 2019 [3]), SciBERT (Beltagy et al., 2019 [4]), and LEGAL-BERT (Chalkidis et al., 2020 [5]) have effectively demonstrated the efficacy of domain-specific embeddings in information retrieval tasks. These endeavors primarily investigated two methodologies: (a) extending the pre-training of BERT and (b) pre-training BERT from scratch, both employing domain-specific corpora. Despite yielding commendable results, these methodologies necessitated substantial domainspecific corpora, with figures as staggering as 21.3B words for BioBERT, 3.17B tokens for SciBERT, and 11.5GB of text data for LEGAL-BERT, thereby posing significant challenges, particularly in low-resource domains like enterprises.
Large Language Models are Biased Reinforcement Learners
Hayes, William M., Yax, Nicolas, Palminteri, Stefano
In-context learning enables large language models (LLMs) to perform a variety of tasks, including learning to make reward-maximizing choices in simple bandit tasks. Given their potential use as (autonomous) decision-making agents, it is important to understand how these models perform such reinforcement learning (RL) tasks and the extent to which they are susceptible to biases. Motivated by the fact that, in humans, it has been widely documented that the value of an outcome depends on how it compares to other local outcomes, the present study focuses on whether similar value encoding biases apply to how LLMs encode rewarding outcomes. Results from experiments with multiple bandit tasks and models show that LLMs exhibit behavioral signatures of a relative value bias. Adding explicit outcome comparisons to the prompt produces opposing effects on performance, enhancing maximization in trained choice sets but impairing generalization to new choice sets. Computational cognitive modeling reveals that LLM behavior is well-described by a simple RL algorithm that incorporates relative values at the outcome encoding stage. Lastly, we present preliminary evidence that the observed biases are not limited to fine-tuned LLMs, and that relative value processing is detectable in the final hidden layer activations of a raw, pretrained model. These findings have important implications for the use of LLMs in decision-making applications.
Human-Generative AI Collaborative Problem Solving Who Leads and How Students Perceive the Interactions
Zhu, Gaoxia, Sudarshan, Vidya, Kow, Jason Fok, Ong, Yew Soon
This research investigates distinct human-generative AI collaboration types and students' interaction experiences when collaborating with generative AI (i.e., ChatGPT) for problem-solving tasks and how these factors relate to students' sense of agency and perceived collaborative problem solving. By analyzing the surveys and reflections of 79 undergraduate students, we identified three human-generative AI collaboration types: even contribution, human leads, and AI leads. Notably, our study shows that 77.21% of students perceived they led or had even contributed to collaborative problem-solving when collaborating with ChatGPT. On the other hand, 15.19% of the human participants indicated that the collaborations were led by ChatGPT, indicating a potential tendency for students to rely on ChatGPT. Furthermore, 67.09% of students perceived their interaction experiences with ChatGPT to be positive or mixed. We also found a positive correlation between positive interaction experience and a sense of positive agency. The results of this study contribute to our understanding of the collaboration between students and generative AI and highlight the need to study further why some students let ChatGPT lead collaborative problem-solving and how to enhance their interaction experience through curriculum and technology design.
HiGPT: Heterogeneous Graph Language Model
Tang, Jiabin, Yang, Yuhao, Wei, Wei, Shi, Lei, Xia, Long, Yin, Dawei, Huang, Chao
Heterogeneous graph learning aims to capture complex relationships and diverse relational semantics among entities in a heterogeneous graph to obtain meaningful representations for nodes and edges. Recent advancements in heterogeneous graph neural networks (HGNNs) have achieved state-of-the-art performance by considering relation heterogeneity and using specialized message functions and aggregation rules. However, existing frameworks for heterogeneous graph learning have limitations in generalizing across diverse heterogeneous graph datasets. Most of these frameworks follow the "pre-train" and "fine-tune" paradigm on the same dataset, which restricts their capacity to adapt to new and unseen data. This raises the question: "Can we generalize heterogeneous graph models to be well-adapted to diverse downstream learning tasks with distribution shifts in both node token sets and relation type heterogeneity?'' To tackle those challenges, we propose HiGPT, a general large graph model with Heterogeneous graph instruction-tuning paradigm. Our framework enables learning from arbitrary heterogeneous graphs without the need for any fine-tuning process from downstream datasets. To handle distribution shifts in heterogeneity, we introduce an in-context heterogeneous graph tokenizer that captures semantic relationships in different heterogeneous graphs, facilitating model adaptation. We incorporate a large corpus of heterogeneity-aware graph instructions into our HiGPT, enabling the model to effectively comprehend complex relation heterogeneity and distinguish between various types of graph tokens. Furthermore, we introduce the Mixture-of-Thought (MoT) instruction augmentation paradigm to mitigate data scarcity by generating diverse and informative instructions. Through comprehensive evaluations, our proposed framework demonstrates exceptional performance in terms of generalization performance.
Cross-Language Assessment of Mathematical Capability of ChatGPT
Sathe, Gargi, Shamraj, Aneesh, Surve, Aditya, Patil, Nahush, Saxena, Kumkum
This paper presents an evaluation of the mathematical capability of ChatGPT across diverse languages like Hindi, Gujarati, and Marathi. ChatGPT, based on GPT-3.5 by OpenAI, has garnered significant attention for its natural language understanding and generation abilities. However, its performance in solving mathematical problems across multiple natural languages remains a comparatively unexplored area, especially in regional Indian languages. In this paper, we explore those capabilities as well as using chain-of-thought prompting to figure out if it increases the accuracy of responses as much as it does in the English language and provide insights into the current limitations.
UrbanGPT: Spatio-Temporal Large Language Models
Li, Zhonghang, Xia, Lianghao, Tang, Jiabin, Xu, Yong, Shi, Lei, Xia, Long, Yin, Dawei, Huang, Chao
Spatio-temporal prediction aims to forecast and gain insights into the ever-changing dynamics of urban environments across both time and space. Its purpose is to anticipate future patterns, trends, and events in diverse facets of urban life, including transportation, population movement, and crime rates. Although numerous efforts have been dedicated to developing neural network techniques for accurate predictions on spatio-temporal data, it is important to note that many of these methods heavily depend on having sufficient labeled data to generate precise spatio-temporal representations. Unfortunately, the issue of data scarcity is pervasive in practical urban sensing scenarios. Consequently, it becomes necessary to build a spatio-temporal model with strong generalization capabilities across diverse spatio-temporal learning scenarios. Taking inspiration from the remarkable achievements of large language models (LLMs), our objective is to create a spatio-temporal LLM that can exhibit exceptional generalization capabilities across a wide range of downstream urban tasks. To achieve this objective, we present the UrbanGPT, which seamlessly integrates a spatio-temporal dependency encoder with the instruction-tuning paradigm. This integration enables LLMs to comprehend the complex inter-dependencies across time and space, facilitating more comprehensive and accurate predictions under data scarcity. To validate the effectiveness of our approach, we conduct extensive experiments on various public datasets, covering different spatio-temporal prediction tasks. The results consistently demonstrate that our UrbanGPT, with its carefully designed architecture, consistently outperforms state-of-the-art baselines. These findings highlight the potential of building large language models for spatio-temporal learning, particularly in zero-shot scenarios where labeled data is scarce.