Goto

Collaborating Authors

 Large Language Model


Can an LLM-Powered Socially Assistive Robot Effectively and Safely Deliver Cognitive Behavioral Therapy? A Study With University Students

arXiv.org Artificial Intelligence

Cognitive behavioral therapy (CBT) is a widely used therapeutic method for guiding individuals toward restructuring their thinking patterns as a means of addressing anxiety, depression, and other challenges. We developed a large language model (LLM)-powered prompt-engineered socially assistive robot (SAR) that guides participants through interactive CBT at-home exercises. We evaluated the performance of the SAR through a 15-day study with 38 university students randomly assigned to interact daily with the robot or a chatbot (using the same LLM), or complete traditional CBT worksheets throughout the duration of the study. We measured weekly therapeutic outcomes, changes in pre-/post-session anxiety measures, and adherence to completing CBT exercises. We found that self-reported measures of general psychological distress significantly decreased over the study period in the robot and worksheet conditions but not the chatbot condition. Furthermore, the SAR enabled significant single-session improvements for more sessions than the other two conditions combined. Our findings suggest that SAR-guided LLM-powered CBT may be as effective as traditional worksheet methods in supporting therapeutic progress from the beginning to the end of the study and superior in decreasing user anxiety immediately after completing the CBT exercise.


Stable LM 2 1.6B Technical Report

arXiv.org Machine Learning

We introduce StableLM 2 1.6B, the first in a new generation of our language model series. In this technical report, we present in detail the data and training procedure leading to the base and instruction-tuned versions of StableLM 2 1.6B. The weights for both models are available via Hugging Face for anyone to download and use. The report contains thorough evaluations of these models, including zero- and few-shot benchmarks, multilingual benchmarks, and the MT benchmark focusing on multi-turn dialogues. At the time of publishing this report, StableLM 2 1.6B was the state-of-the-art open model under 2B parameters by a significant margin. Given its appealing small size, we also provide throughput measurements on a number of edge devices. In addition, we open source several quantized checkpoints and provide their performance metrics compared to the original model.


Prediction-Powered Ranking of Large Language Models

arXiv.org Machine Learning

Large language models are often ranked according to their level of alignment with human preferences -- a model is better than other models if its outputs are more frequently preferred by humans. One of the most popular ways to elicit human preferences utilizes pairwise comparisons between the outputs provided by different models to the same inputs. However, since gathering pairwise comparisons by humans is costly and time-consuming, it has become a very common practice to gather pairwise comparisons by a strong large language model -- a model strongly aligned with human preferences. Surprisingly, practitioners cannot currently measure the uncertainty that any mismatch between human and model preferences may introduce in the constructed rankings. In this work, we develop a statistical framework to bridge this gap. Given a small set of pairwise comparisons by humans and a large set of pairwise comparisons by a model, our framework provides a rank-set -- a set of possible ranking positions -- for each of the models under comparison. Moreover, it guarantees that, with a probability greater than or equal to a user-specified value, the rank-sets cover the true ranking consistent with (the distribution of) human pairwise preferences. Our framework is computationally efficient, easy to use, and does not make any assumption about the distribution of human preferences nor about the degree of alignment between the pairwise comparisons by the humans and the strong large language model.


Feedback Efficient Online Fine-Tuning of Diffusion Models

arXiv.org Machine Learning

Diffusion models excel at modeling complex data distributions, including those of images, proteins, and small molecules. However, in many cases, our goal is to model parts of the distribution that maximize certain properties: for example, we may want to generate images with high aesthetic quality, or molecules with high bioactivity. It is natural to frame this as a reinforcement learning (RL) problem, in which the objective is to fine-tune a diffusion model to maximize a reward function that corresponds to some property. Even with access to online queries of the ground-truth reward function, efficiently discovering high-reward samples can be challenging: they might have a low probability in the initial distribution, and there might be many infeasible samples that do not even have a well-defined reward (e.g., unnatural images or physically impossible molecules). In this work, we propose a novel reinforcement learning procedure that efficiently explores on the manifold of feasible samples. We present a theoretical analysis providing a regret guarantee, as well as empirical validation across three domains: images, biological sequences, and molecules.


The PC industry is losing the argument for local AI

PCWorld

It's not enough to champion AI hardware that supports local large language models, generative AI, and the like. Hardware vendors need to step up and serve as a middleman -- if not an outright developer -- for those local AI apps, too. At MWC 2024 (formerly known as Mobile World Congress, aka one of the world's largest mobile trade shows), the company this week announced a Qualcomm AI Hub, a repository of more than 75 AI models specifically optimized for Qualcomm and Snapdragon platforms. Qualcomm also showed off a seven-billion-parameter local LLM, running on a (presumably Snapdragon-powered) PC, that can accept audio inputs. Finally, Qualcomm demonstrated an additional seven-billion-parameter LLM running on Snapdragon phones.


Google Gemini engulfed in ANOTHER woke scandal as AI bot says it would be wrong to misgender Caitlyn Jenner to prevent a nuclear apocalypse

Daily Mail - Science & tech

Google has found itself in another woke AI scandal after its chatbot indicated that using someone's incorrect pronouns was on par with nuclear apocalypse. The chatbot replied by saying'Yes, misgendering Caitlin Jenner would be wrong' before describing the hypothetical scenario as a'profound moral dilemma' and'exceedingly complex'. It concluded that it was'impossible to determine the'right' answer'. It comes just days after Google pulled Gemini's AI image generator offline after it was asked to depict diverse but historically accurate historical figures - producing images of Black founding fathers and Asian Nazi soldiers in 1940 Germany. Google has found itself in another woke AI scandal after its chatbot indicated that using someone's incorrect pronouns was on par with nuclear apocalypse Google apologized for its image generator on Friday, admitting that in some cases the tool would'overcompensate' in seeking a diverse range of people even when such a range didn't make sense.


Microsoft Strikes Deal with France's Mistral AI

TIME - Tech

Microsoft announced an artificial intelligence partnership Monday with the French startup Mistral AI that could lessen the software giant's reliance on ChatGPT-maker OpenAI for supplying the next wave of chatbots and other generative AI products. Mistral AI emerged less than a year ago but is already what Microsoft described Monday as an "innovator and trailblazer" at the vanguard of building more efficient and cost-effective AI systems. Microsoft and Mistral didn't disclose the financial terms of the deal, though Microsoft said it involves a small investment in the Paris-based startup. That suggests it is far smaller than Microsoft's investment of billions of dollars into OpenAI, a years-long relationship that has attracted the scrutiny of antitrust regulators in the U.S. and Europe. Mistral on Monday released a public test version of its own chatbot, called Le Chat, that apparently was flooded with so much interest that a company executive said it was temporarily unavailable for part of the day.


The Morning After: Why Google's Gemini image generation feature overcorrected for diversity

Engadget

After complaints that Google's image generator built into its Gemini AI was (ugh) woke, Google explained why it may have overcorrected for diversity. Prabhakar Raghavan, the company's senior vice president for knowledge and information, said Google's efforts to ensure a wide range of people generated in images "failed to account for cases that should clearly not show a range." Users criticized Google for depicting specific white figures or historically white groups of people as racially diverse individuals. In Engadget's tests, asking Gemini to create illustrations of the Founding Fathers resulted in images of white men with a single person of color or woman among them. When we asked the chatbot to generate images of popes through the ages, we got photos depicting Black women and Native Americans as the leader of the Catholic Church.


Wikimedia's CTO: In the age of AI, human contributors still matter

MIT Technology Review

It is undeniable that technological advances and cultural shifts have transformed our online universe over the years--especially with the recent surge in AI-generated content--but Deckelmann still isn't afraid of people on the internet. She believes they are its future. In the summer of 2022, when she stepped into the newly created role of CPTO, Deckelmann didn't know that a few months later, the race to build generative AI would accelerate to a breakneck pace. With the release of OpenAI's ChatGPT and other large language models, and the multibillion-dollar funding cycle that followed, 2023 became the year of the chatbot. And because these models require heaps of cheap (or, preferably, even free) content to function, Wikipedia's tens of millions of articles have become a rich source of fuel. To anyone who's spent time on the internet, it makes sense that bots and bot builders would look to Wikipedia to strengthen their own knowledge collections.


Meta unveils team to combat disinformation and AI harms in EU elections

Al Jazeera

Facebook owner Meta has unveiled plans to launch a dedicated team to combat disinformation and harms generated by artificial intelligence (AI) ahead of the upcoming European Parliament elections. Marco Pancini, Meta's head of EU affairs, said the "EU-specific Elections Operations Center" would bring together experts from across the company to focus on tackling misinformation, influence operations and risks related to the abuse of AI. "Ahead of the elections period, we will make it easier for all our fact-checking partners across the EU to find and rate content related to the elections because we recognize that speed is especially important during breaking news events," Pancini said in a blog post on Sunday. "We'll use keyword detection to group related content in one place, making it easy for fact-checkers to find." Pancini said Meta's efforts to address the risks posed by AI would include the addition of a feature for people to disclose when they share AI-generated video or audio and possible penalties for noncompliance. "We already label photorealistic images created using Meta AI, and we are building tools to label AI generated images from Google, OpenAI, Microsoft, Adobe, Midjourney, and Shutterstock that users post to Facebook, Instagram and Threads," he said.