Goto

Collaborating Authors

 Large Language Model


Data Augmentations for Improved (Large) Language Model Generalization

arXiv.org Artificial Intelligence

The reliance of text classifiers on spurious correlations can lead to poor generalization at deployment, raising concerns about their use in safety-critical domains such as healthcare. In this work, we propose to use counterfactual data augmentation, guided by knowledge of the causal structure of the data, to simulate interventions on spurious features and to learn more robust text classifiers. We show that this strategy is appropriate in prediction problems where the label is spuriously correlated with an attribute. Under the assumptions of such problems, we discuss the favorable sample complexity of counterfactual data augmentation, compared to importance re-weighting. Pragmatically, we match examples using auxiliary data, based on diff-in-diff methodology, and use a large language model (LLM) to represent a conditional probability of text. Through extensive experimentation on learning caregiver-invariant predictors of clinical diagnoses from medical narratives and on semi-synthetic data, we demonstrate that our method for simulating interventions improves out-of-distribution (OOD) accuracy compared to baseline invariant learning algorithms.


Bias Testing and Mitigation in LLM-based Code Generation

arXiv.org Artificial Intelligence

Utilizing state-of-the-art Large Language Models (LLMs), automatic code generation models play a pivotal role in enhancing the productivity of software development procedures. As the adoption of LLMs becomes more widespread in software coding ecosystems, a pressing issue has emerged: does the generated code contain social bias and unfairness, such as those related to age, gender, and race? This issue concerns the integrity, fairness, and ethical foundation of software applications that depend on the code generated by these models, yet is under-explored in the literature. This paper presents a novel bias testing framework that is specifically designed for code generation tasks. Based on this framework, we conduct an extensive evaluation of the bias in code generated by five state-of-the-art LLMs. Our findings reveal that 20.29% to 44.93% code functions generated by the models under study are biased when handling bias sensitive tasks (i.e., tasks that involve sensitive attributes such as age and gender). This indicates that the existing LLMs can be unfair in code generation, posing risks of unintended and harmful software behaviors. To mitigate bias for code generation models, we evaluate five bias mitigation prompt strategies, i.e., utilizing bias testing results to refine the code (zero-shot), one-, few-shot, and two Chain-of-Thought (CoT) prompts. Our evaluation results illustrate that these strategies are all effective in mitigating bias. Overall, one-shot and few-shot learning are the two most effective. For GPT-4, 80% to 90% code bias can be removed with one-shot learning.


Where Would I Go Next? Large Language Models as Human Mobility Predictors

arXiv.org Artificial Intelligence

Accurate human mobility prediction underpins many important applications across a variety of domains, including epidemic modelling, transport planning, and emergency responses. Due to the sparsity of mobility data and the stochastic nature of people's daily activities, achieving precise predictions of people's locations remains a challenge. While recently developed large language models (LLMs) have demonstrated superior performance across numerous language-related tasks, their applicability to human mobility studies remains unexplored. Addressing this gap, this article delves into the potential of LLMs for human mobility prediction tasks. We introduce a novel method, LLM-Mob, which leverages the language understanding and reasoning capabilities of LLMs for analysing human mobility data. We present concepts of historical stays and context stays to capture both long-term and short-term dependencies in human movement and enable time-aware prediction by using time information of the prediction target. Additionally, we design context-inclusive prompts that enable LLMs to generate more accurate predictions. Comprehensive evaluations of our method reveal that LLM-Mob excels in providing accurate and interpretable predictions, highlighting the untapped potential of LLMs in advancing human mobility prediction techniques. We posit that our research marks a significant paradigm shift in human mobility modelling, transitioning from building complex domain-specific models to harnessing general-purpose LLMs that yield accurate predictions through language instructions. The code for this work is available at https://github.com/xlwang233/LLM-Mob.


The Unequal Opportunities of Large Language Models: Revealing Demographic Bias through Job Recommendations

arXiv.org Artificial Intelligence

Large Language Models (LLMs) have seen widespread deployment in various real-world applications. Understanding these biases is crucial to comprehend the potential downstream consequences when using LLMs to make decisions, particularly for historically disadvantaged groups. In this work, we propose a simple method for analyzing and comparing demographic bias in LLMs, through the lens of job recommendations. We demonstrate the effectiveness of our method by measuring intersectional biases within ChatGPT and LLaMA, two cutting-edge LLMs. Our experiments primarily focus on uncovering gender identity and nationality bias; however, our method can be extended to examine biases associated with any intersection of demographic identities. We identify distinct biases in both models toward various demographic identities, such as both models consistently suggesting low-paying jobs for Mexican workers or preferring to recommend secretarial roles to women. Our study highlights the importance of measuring the bias of LLMs in downstream applications to understand the potential for harm and inequitable outcomes.


Getting ViT in Shape: Scaling Laws for Compute-Optimal Model Design

arXiv.org Artificial Intelligence

Scaling laws have been recently employed to derive compute-optimal model size (number of parameters) for a given compute duration. We advance and refine such methods to infer compute-optimal model shapes, such as width and depth, and successfully implement this in vision transformers. Our shape-optimized vision transformer, SoViT, achieves results competitive with models that exceed twice its size, despite being pre-trained with an equivalent amount of compute. For example, SoViT-400m/14 achieves 90.3% fine-tuning accuracy on ILSRCV2012, surpassing the much larger ViT-g/14 and approaching ViT-G/14 under identical settings, with also less than half the inference cost. We conduct a thorough evaluation across multiple tasks, such as image classification, captioning, VQA and zero-shot transfer, demonstrating the effectiveness of our model across a broad range of domains and identifying limitations. Overall, our findings challenge the prevailing approach of blindly scaling up vision models and pave a path for a more informed scaling.


LaMP: When Large Language Models Meet Personalization

arXiv.org Artificial Intelligence

This paper highlights the importance of personalization in large language models and introduces the LaMP benchmark -- a novel benchmark for training and evaluating language models for producing personalized outputs. LaMP offers a comprehensive evaluation framework with diverse language tasks and multiple entries for each user profile. It consists of seven personalized tasks, spanning three text classification and four text generation tasks. We additionally propose two retrieval augmentation approaches that retrieve personal items from each user profile for personalizing language model outputs. To this aim, we study various retrieval models, including term matching, semantic matching, and time-aware methods. Extensive experiments on LaMP for zero-shot and fine-tuned language models demonstrate the efficacy of the proposed retrieval augmentation approach and highlight the impact of personalization in various natural language tasks.


Volkswagen thinks ChatGPT integration will make its in-car voice assistant good

Engadget

AI is literally everywhere, so it's not a big surprise to learn that Volkswagen is planning to bring ChatGPT to its vehicles. As part of its CES 2024 announcements, the automaker says that its existing IDA voice assistant will work with ChatGPT across a range of its newer models. VW isn't the first to try this -- Mercedes-Benz announced ChatGPT integration in June of last year, so it seems like this is certainly a thing we're all going to have to get used to. Specifically, VW says that ChatGPT will be enabled in these specific models with the latest generation of the company's infotainment systems: ID.7 (pictured above), ID.4, It'll roll out ChatGPT as as "standard feature" in "many" production vehicles in Q2 of 2024; the company didn't say in which regions, but notes that the feature is only currently "being considered" for the US market.


AMD's Ryzen 8000 brings AI to the desktop, with an AM4 surprise

PCWorld

Turnabout is fair play: At CES 2024, AMD launched desktop versions of its mobile Ryzen 8000 processors, bringing AI to the desktop alongside integrated graphics. And for AMD fans who aren't quite ready to make the leap to AMD's AM5 socket, whoa! There are new Ryzen 5000 desktop chips as well. AMD's launch adds four new Ryzen 8000 G-series processors to AMD's lineup. These are APUs, AMD's desktop chips that combine integrated graphics alongside the CPU die -- in this case, the RDNA 3- based Radeon 780M, Radeon 760M, and Radeon 740M that we saw integrated in the AMD Ryzen 8040 (8000) series chips AMD announced this December. All of the new Ryzen 8000 and 5000 chips will be available on Jan. 31.


'Impossible' to create AI tools like ChatGPT without copyrighted material, OpenAI says

The Guardian

Last month, the New York Times sued OpenAI and Microsoft, which is a leading investor in OpenAI and uses its tools in its products, accusing them of "unlawful use" of its work to create their products. Responding to the NYT lawsuit last month, OpenAI had said it respected "the rights of content creators and owners". The NYT lawsuit has followed numerous other legal complaints against OpenAI. John Grisham, Jodi Picoult and George RR Martin were among 17 authors who sued OpenAI in September alleging "systematic theft on a mass scale". Get set for the working day – we'll point you to all the business news and analysis you need every morning Elsewhere in its House of Lords submission, in response to a question about AI safety, OpenAI said it supported independent analysis of its security measures.


Four lessons from 2023 that tell us where AI regulation is going

MIT Technology Review

Most broadly, we are likely to see the strategies that emerged last year continue, expand, and begin to be implemented. For example, following President Biden's executive order, various US government agencies may outline new best practices but empower AI companies to police themselves. And across the pond, companies and regulators will begin to grapple with Europe's AI Act and its risk-based approach. It certainly won't be seamless, and there's bound to be a lot of discussion about how these new laws and policies actually work in practice. While writing this piece, I took some time to reflect on how we got here.