Deep Learning
Bayesian Evaluation of Large Language Model Behavior
Longjohn, Rachel, Wu, Shang, Kher, Saatvik, Belรฉm, Catarina, Smyth, Padhraic
It is increasingly important to evaluate how text generation systems based on large language models (LLMs) behave, such as their tendency to produce harmful output or their sensitivity to adversarial inputs. Such evaluations often rely on a curated benchmark set of input prompts provided to the LLM, where the output for each prompt may be assessed in a binary fashion (e.g., harmful/non-harmful or does not leak/leaks sensitive information), and the aggregation of binary scores is used to evaluate the LLM. However, existing approaches to evaluation often neglect statistical uncertainty quantification. With an applied statistics audience in mind, we provide background on LLM text generation and evaluation, and then describe a Bayesian approach for quantifying uncertainty in binary evaluation metrics. We focus in particular on uncertainty that is induced by the probabilistic text generation strategies typically deployed in LLM-based systems. We present two case studies applying this approach: 1) evaluating refusal rates on a benchmark of adversarial inputs designed to elicit harmful responses, and 2) evaluating pairwise preferences of one LLM over another on a benchmark of open-ended interactive dialogue examples. We demonstrate how the Bayesian approach can provide useful uncertainty quantification about the behavior of LLM-based systems.
How Google's DeepMind tool is 'more quickly' forecasting hurricane behavior
How Google's DeepMind tool is'more quickly' forecasting hurricane behavior'Less expensive and time consuming' model helps with fast and accurate predictions, possibly saving lives and property When then Tropical Storm Melissa was churning south of Haiti, Philippe Papin, a National Hurricane Center (NHC) meteorologist, had confidence it was about to grow into a monster hurricane. As the lead forecaster on duty, he predicted that in just 24 hours the storm would become a category 4 hurricane and begin a turn towards the coast of Jamaica. No NHC forecaster had ever issued such a bold forecast for rapid strengthening. But Papin had an ace up his sleeve: artificial intelligence in the form of Google's new DeepMind hurricane model - released for the first time in June. And, as predicted, Melissa did become a storm of astonishing strength that tore through Jamaica.
Use Google Gemini and ChatGPT to Organize Your Life With Scheduled Actions
The AI's latest trick is following the schedule you set for it. The developers of the big generative AI chatbots are continuing to push out new features at a rapid rate, as they bid to make sure their bot is the one you turn to whenever you need some assistance from artificial intelligence. One of the latest updates to Google Gemini gives you the ability to set up scheduled actions. These are exactly what they sound like: Tasks that you can get Google Gemini to run automatically, on a schedule. Maybe you want a weather and news report every morning at 7 am, or perhaps you want an evening meal suggestion every evening at 7 pm.
Meet the all-in-one AI platform that could replace every tool you use, now 86% off
When you purchase through links in our articles, we may earn a small commission. For $74.97 (MSRP $540), get a lifetime subscription to a 1minAI Advanced Business Plan, an all-in-one platform that uses multiple top AI models. You know ChatGPT, Gemini, and all the usual heavy-hitters -- but you've probably missed the newcomer quickly climbing the ranks. Get permanent access to an Advanced Business Plan on sale for $74.97 for a limited time. So, you need to generate articles for work?