Large Language Model
Harmful Speech Detection by Language Models Exhibits Gender-Queer Dialect Bias
Dorn, Rebecca, Kezar, Lee, Morstatter, Fred, Lerman, Kristina
Content moderation on social media platforms shapes the dynamics of online discourse, influencing whose voices are amplified and whose are suppressed. Recent studies have raised concerns about the fairness of content moderation practices, particularly for aggressively flagging posts from transgender and non-binary individuals as toxic. In this study, we investigate the presence of bias in harmful speech classification of gender-queer dialect online, focusing specifically on the treatment of reclaimed slurs. We introduce a novel dataset, QueerReclaimLex, based on 109 curated templates exemplifying non-derogatory uses of LGBTQ+ slurs. Dataset instances are scored by gender-queer annotators for potential harm depending on additional context about speaker identity. We systematically evaluate the performance of five off-the-shelf language models in assessing the harm of these texts and explore the effectiveness of chain-of-thought prompting to teach large language models (LLMs) to leverage author identity context. We reveal a tendency for these models to inaccurately flag texts authored by gender-queer individuals as harmful. Strikingly, across all LLMs the performance is poorest for texts that show signs of being written by individuals targeted by the featured slur (F1 <= 0.24). We highlight an urgent need for fairness and inclusivity in content moderation systems. By uncovering these biases, this work aims to inform the development of more equitable content moderation practices and contribute to the creation of inclusive online spaces for all users.
On the Worst Prompt Performance of Large Language Models
Cao, Bowen, Cai, Deng, Zhang, Zhisong, Zou, Yuexian, Lam, Wai
The performance of large language models (LLMs) is acutely sensitive to the phrasing of prompts, which raises significant concerns about their reliability in real-world scenarios. Existing studies often divide prompts into task-level instructions and case-level inputs and primarily focus on evaluating and improving robustness against variations in tasks-level instructions. However, this setup fails to fully address the diversity of real-world user queries and assumes the existence of task-specific datasets. To address these limitations, we introduce RobustAlpacaEval, a new benchmark that consists of semantically equivalent case-level queries and emphasizes the importance of using the worst prompt performance to gauge the lower bound of model performance. Extensive experiments on RobustAlpacaEval with ChatGPT and six open-source LLMs from the Llama, Mistral, and Gemma families uncover substantial variability in model performance; for instance, a difference of 45.48% between the worst and best performance for the Llama-2-70B-chat model, with its worst performance dipping as low as 9.38%. We further illustrate the difficulty in identifying the worst prompt from both model-agnostic and model-dependent perspectives, emphasizing the absence of a shortcut to characterize the worst prompt. We also attempt to enhance the worst prompt performance using existing prompt engineering and prompt consistency methods, but find that their impact is limited. These findings underscore the need to create more resilient LLMs that can maintain high performance across diverse prompts. Data and code are available at https://github.com/cbwbuaa/On-the-Worst-Prompt- Performance-of-LLMs.
Battling Botpoop using GenAI for Higher Education: A Study of a Retrieval Augmented Generation Chatbots Impact on Learning
Thway, Maung, Recatala-Gomez, Jose, Lim, Fun Siong, Hippalgaonkar, Kedar, Ng, Leonard W. T.
Generative artificial intelligence (GenAI) and large language models (LLMs) have simultaneously opened new avenues for enhancing human learning and increased the prevalence of poor-quality information in student response - termed'Botpoop'. This study introduces Professor Leodar, a custom-built, Singlish-speaking Retrieval Augmented Generation (RAG) chatbot designed to enhance educational while reducing Botpoop. Deployed at Nanyang Technological University, Singapore, Professor Leodar offers a glimpse into the future of AI-assisted learning, offering personalized guidance, 24/7 availability, and contextually relevant information. Through a mixed-methods approach, we examine the impact of Professor Leodar on learning, engagement, and exam preparedness, with 97.1% of participants reporting positive experiences. These findings help define possible roles of AI in education and highlight the potential of custom GenAI chatbots. Our combination of chatbot development, in-class deployment and outcomes study offers a benchmark for GenAI educational tools and is a stepping stone for redefining the interplay between AI and human learning.
SAIL: Self-Improving Efficient Online Alignment of Large Language Models
Ding, Mucong, Chakraborty, Souradip, Agrawal, Vibhu, Che, Zora, Koppel, Alec, Wang, Mengdi, Bedi, Amrit, Huang, Furong
As artificial intelligence (AI) systems surpass human capabilities in various tasks, ensuring alignment with human values and ethics is crucial. This is especially important for large language models (LLMs), which are trained on diverse datasets that may contain harmful content. Reinforcement Learning from Human Feedback (RLHF) is an effective method for AI alignment, with models like OpenAI's GPT-4, Google's Gemini, and Anthropic Claude showing safe and aligned behaviors. However, the vast majority of the current research in RLHF (Agarwal et al., 2020; Rafailov et al., 2023; Ouyang et al., 2022; Chakraborty et al., 2024; Swamy et al., 2024) focuses on the offline setting, which involves using a fixed dataset of responses generated by the supervised fine-tuned model (SFT), ranked by human experts. Consequently, these methods are inherently offline and heavily reliant on the quality of the offline data generated by the SFT model, which exhibits drawbacks such as insufficient coverage of response-query pairs leading to sub-optimal alignment. To deal with the above shortcomings, recent literature (Guo et al., 2024a; Sharma et al., 2024; Lee et al., 2023; Yuan et al., 2024b) has focused on designing online RLHF algorithms. The setting of online RLHF transcends the constraints of a static offline dataset and aims to address two critical questions: Q1: How should we generate new responses during fine-tuning?
Anthropic Touts New AI Model as 'Most Intelligent Yet'
Anthropic launched a new AI model Thursday which it says is its "most intelligent model yet." The new model, Claude 3.5 Sonnet, is reportedly twice as fast as Claude 3 Opus, the company's previous best-in-class AI, and five times cheaper to run. Following a trend set by its competitor OpenAI--which just last month released the newest version of ChatGPT, GPT-4o--Claude 3.5 Sonnet is free for all users on the web and iOS to access. It has also been made available to developers. Claude 3.5 Sonnet is "now the most intelligent model in the world," claims Michael Gerstenhaber, a product manager at the company.
Anthropic's newest Claude chatbot beats OpenAI's GPT-4o in some benchmarks
Anthropic rolled out its newest AI language model on Thursday, Claude 3.5 Sonnet. The updated chatbot outperforms the company's previous top-tier model, Claude 3 Opus, while working at twice the speed. Claude users (including those on free accounts) can check it out beginning today. Sonnet, which tends to be Anthropic's most balanced model, is the first release in the Claude 3.5 family. The company says Claude 3.5 Haiku (the fastest in each generation) and Claude 3.5 Opus (the most powerful) will arrive later this year.
Good news! Most apps I've tried on Microsoft's Copilot Surface just work
Have you been burned before by Windows on Arm? Are you worried whether the apps you need will actually run on Copilot PCs? But after playing around with one myself, I'm fairly optimistic that those days are over, as Qualcomm executives promised. After receiving a Surface Pro (2024) 11th Edition from Microsoft for review, I spent a good chunk of my first day just downloading various applications and seeing if they'd run--and if they did, how well. First, this is indeed a productivity tablet, and Microsoft and Qualcomm have done a good job making sure most common productivity application work without hassle. Second, Copilot PCs are not gaming PCs, and there's a good chance your favorite games won't even run.
We're Still Waiting for the Next Big Leap in AI
When OpenAI announced GPT-4, its latest large language model, last March, it sent shockwaves through the tech world. It was clearly more capable than anything seen before at chatting, coding, and solving all sorts of thorny problems--including school homework. Anthropic, a rival to OpenAI, announced today that it has made its own AI advance that will upgrade chatbots and other use cases. But although the new model is the world's best by some measures, it's more of a step forward than a big leap. Anthropic's new model, called Claude 3.5 Sonnet, is an upgrade to its existing Claude 3 family of AI models.
Europe Scrambles for Relevance in the Age of AI
When a Finn talks to an AI helper like ChatGPT, they often get the sense that something is subtly wrong. "You really feel that this conversation is not the way that you would have a discussion in Finland," says Peter Sarlin. For a start, Finnish people are known for a blunt approach to dialog and chatbots are usually tuned to be overly courteous. But there's also the fact that most leading chatbots and the large language models behind them are developed in the US and trained on mostly US data. Cutting-edge AI products often come with a tonality that is essentially American.
ReaLHF: Optimized RLHF Training for Large Language Models through Parameter Reallocation
Mei, Zhiyu, Fu, Wei, Li, Kaiwei, Wang, Guangju, Zhang, Huanchen, Wu, Yi
Reinforcement Learning from Human Feedback (RLHF) stands as a pivotal technique in empowering large language model (LLM) applications. Since RLHF involves diverse computational workloads and intricate dependencies among multiple LLMs, directly adopting parallelization techniques from supervised training can result in sub-optimal performance. To overcome this limitation, we propose a novel approach named parameter ReaLlocation, which dynamically redistributes LLM parameters in the cluster and adapts parallelization strategies during training. Building upon this idea, we introduce ReaLHF, a pioneering system capable of automatically discovering and running efficient execution plans for RLHF training given the desired algorithmic and hardware configurations. ReaLHF formulates the execution plan for RLHF as an augmented dataflow graph. Based on this formulation, ReaLHF employs a tailored search algorithm with a lightweight cost estimator to discover an efficient execution plan. Subsequently, the runtime engine deploys the selected plan by effectively parallelizing computations and redistributing parameters. We evaluate ReaLHF on the LLaMA-2 models with up to $4\times70$ billion parameters and 128 GPUs. The experiment results showcase ReaLHF's substantial speedups of $2.0-10.6\times$ compared to baselines. Furthermore, the execution plans generated by ReaLHF exhibit an average of $26\%$ performance improvement over heuristic approaches based on Megatron-LM. The source code of ReaLHF is publicly available at https://github.com/openpsi-project/ReaLHF .