Large Language Model
Hallucination Detox: Sensitivity Dropout (SenD) for Large Language Model Training
Mohammadzadeh, Shahrad, Guerra, Juan David, Bonizzato, Marco, Rabbany, Reihaneh, Farnadi, Golnoosh
As large language models (LLMs) are increasingly deployed across various industries, concerns regarding their reliability, particularly due to hallucinations - outputs that are factually inaccurate or irrelevant to user input - have grown. Our research investigates the relationship between the training process and the emergence of hallucinations to address a key gap in existing research that focuses primarily on post hoc detection and mitigation strategies. Using models from the Pythia suite (70M - 12B parameters) and several hallucination detection metrics, we analyze hallucination trends throughout training and explore LLM internal dynamics. We introduce Sensitivity Dropout (SenD), a novel training protocol designed to mitigate hallucinations by reducing variance during training. SenD achieves this by deterministically dropping embedding indices with significant variability, referred to as Sensitive Embedding Indices. In addition, we develop an unsupervised hallucination detection metric, Efficient EigenScore (EES), which approximates the traditional EigenScore at 2x speed. This efficient metric is integrated into our protocol, allowing SenD to be both computationally scalable and effective at reducing hallucinations. Our empirical evaluation demonstrates that our approach improves LLM reliability at test time by up to 40% compared to normal training while also providing an efficient method to improve factual accuracy when adapting LLMs to Wikipedia, Medical, and LegalBench domains.
Rescriber: Smaller-LLM-Powered User-Led Data Minimization for Navigating Privacy Trade-offs in LLM-Based Conversational Agent
Zhou, Jijie, Xu, Eryue, Wu, Yaoyao, Li, Tianshi
The proliferation of LLM-based conversational agents has resulted in excessive disclosure of identifiable or sensitive information. However, existing technologies fail to offer perceptible control or account for users' personal preferences about privacy-utility tradeoffs due to the lack of user involvement. To bridge this gap, we designed, built, and evaluated Rescriber, a browser extension that supports user-led data minimization in LLM-based conversational agents by helping users detect and sanitize personal information in their prompts. Our studies (N=12) showed that Rescriber helped users reduce unnecessary disclosure and addressed their privacy concerns. Users' subjective perceptions of the system powered by Llama3-8B were on par with that by GPT-4o. The comprehensiveness and consistency of the detection and sanitization emerge as essential factors that affect users' trust and perceived protection. Our findings confirm the viability of smaller-LLM-powered, user-facing, on-device privacy controls, presenting a promising approach to address the privacy and trust challenges of AI.
MMAD: The First-Ever Comprehensive Benchmark for Multimodal Large Language Models in Industrial Anomaly Detection
Jiang, Xi, Li, Jian, Deng, Hanqiu, Liu, Yong, Gao, Bin-Bin, Zhou, Yifeng, Li, Jialin, Wang, Chengjie, Zheng, Feng
In the field of industrial inspection, Multimodal Large Language Models (MLLMs) have a high potential to renew the paradigms in practical applications due to their robust language capabilities and generalization abilities. However, despite their impressive problem-solving skills in many domains, MLLMs' ability in industrial anomaly detection has not been systematically studied. To bridge this gap, we present MMAD, the first-ever full-spectrum MLLMs benchmark in industrial Anomaly Detection. We defined seven key subtasks of MLLMs in industrial inspection and designed a novel pipeline to generate the MMAD dataset with 39,672 questions for 8,366 industrial images. With MMAD, we have conducted a comprehensive, quantitative evaluation of various state-of-theart MLLMs. The commercial models performed the best, with the average accuracy of GPT-4o models reaching 74.9%. However, this result falls far short of industrial requirements. Our analysis reveals that current MLLMs still have significant room for improvement in answering questions related to industrial anomalies and defects. We further explore two training-free performance enhancement strategies to help models improve in industrial scenarios, highlighting their promising potential for future research. The code and data are available at https://github.com/jam-cc/MMAD. Automatic vision inspection is a crucial challenge in realizing an unmanned factory (Benbarrad et al., 2021). Traditional AI research for automatic vision inspection, such as industrial anomaly detection (IAD) (Jiang et al., 2022b; Ren et al., 2022), typically relies on discriminative models within the conventional deep learning paradigm. These models can only perform trained detection tasks and cannot provide detailed reports like quality inspection workers. The development of MLLMs (Jin et al., 2024) has the potential to alter this situation. These generative models can flexibly produce the required textual output based on input language and visual prompts, allowing us to guide the model using language similar to instructing humans. Nowadays, multimodal large language models, represented by GPT-4 (Achiam et al., 2023), can already do many human jobs, especially high-paying intellectual jobs like programmers, writers, and data analysts (Eloundou et al., 2023). In comparison, the work of quality inspectors is simple, typically not requiring a high level of education but relying heavily on work experience.
Deploying Open-Source Large Language Models: A performance Analysis
Bendi-Ouis, Yannis, Dutartre, Dan, Hinaut, Xavier
Since the release of ChatGPT in November 2022, large language models (LLMs) have seen considerable success, including in the open-source community, with many open-weight models available. However, the requirements to deploy such a service are often unknown and difficult to evaluate in advance. To facilitate this process, we conducted numerous tests at the Centre Inria de l'Universit\'e de Bordeaux. In this article, we propose a comparison of the performance of several models of different sizes (mainly Mistral and LLaMa) depending on the available GPUs, using vLLM, a Python library designed to optimize the inference of these models. Our results provide valuable information for private and public groups wishing to deploy LLMs, allowing them to evaluate the performance of different models based on their available hardware. This study thus contributes to facilitating the adoption and use of these large language models in various application domains.
AMD's Ryzen AI Max is a one-of-a-kind graphics and AI powerhouse
AMD has launched what one executive called "the most advanced mobile X86 processor ever created" at CES 2025: The Ryzen AI Max and AI Max, with absolutely massive capabilities to run graphics and AI workloads. AMD is positioning this "Strix Halo" chip as a sort of hybrid for graphics and AI workstations, comparing it to Nvidia's existing GeForce 4090 GPU in terms of running AI LLMs at up to 70 billion parameters. But the Ryzen AI Max offers more than just that. It's an APU with graphics capabilities that push into discrete GPU territory -- in the 3DMark Steel Nomad benchmark, for example, the chip offers 258 percent the graphics performance of Intel's Core Ultra 9 288V (Arrow Lake) CPU. It also offers excellent AI performance, both with or without the NPU.
Microsoft wants to run Copilot locally on your PC starting early 2025
Microsoft will bring Phi Silica to the Windows runtime this quarter as part of Copilot, according to Microsoft's head of Windows devices, who presented on Monday at CES 2025 in Las Vegas. Microsoft debuted Phi Silica at its Build conference in Seattle last May, showing off the Small Language Model (SLM) that's meant to complement its Large Language Model (LLM) that runs in the cloud. Phi Silica paves the way for a local version of Copilot to run on Windows PCs. Typically, LLMs are faster and more accurate than SLMs. However, they need to run in the cloud and can require expensive subscriptions for full access. On the other hand, SLMs can run AI chatbots and other AI-driven applications on a local PC, but they're less sophisticated and they require NPUs that provide local AI capabilities for PCs, which can ensure privacy and prevent information from leaking to the cloud.
'Virtual employees' could join workforce as soon as this year, OpenAI boss says
Virtual employees could join workforces this year and transform how companies work, according to the chief executive of OpenAI. The first artificial intelligence agents may start working for organisations this year, wrote Sam Altman, as AI firms push for uses that generate returns on substantial investment in the technology. Microsoft, the biggest backer of the company behind ChatGPT, has already announced the introduction of AI agents – tools that can carry out tasks autonomously – with the blue-chip consulting firm McKinsey among the early adopters. "We believe that, in 2025, we may see the first AI agents'join the workforce' and materially change the output of companies," wrote Altman in a blogpost published on Monday. OpenAI is reportedly planning to launch an AI agent codenamed "Operator" this month, after Microsoft announced its Copilot Studio product and rival Anthropic launched the Claude 3.5 Sonnet AI model, which can carry out tasks on the computer such as moving a mouse cursor and typing text.
AI means the end of internet search as we've known it
The biggest change to the way search engines have delivered information to us since the 1990s is happening right now. Which means instead of keywords, you use real questions, expressed in natural language. And instead of links, you'll increasingly be met with answers, written by generative AI and based on live information from all across the internet, delivered the same way. Of course, Google--the company that has defined search for the past 25 years--is trying to be out front on this. In May of 2023, it began testing AI-generated responses to search queries, using its large language model (LLM) to deliver the kinds of answers you might expect from an expert source or trusted friend.
LG Gram Pro 2-in-1 (2025) hands-on: Of course a thin and light laptop gets AI at CES 2025
It's been ten years since LG introduced its Gram line of ultra thin and light laptops, and despite my early skepticism about its longevity and build quality, the company continues to make new models regularly. It's expanded the portfolio to offer pro variants, clamshells and 2-in-1s, and in keeping with every laptop maker in recent years, LG is now infusing the Gram Pros with more of its own... you guessed it... AI. We already learned about this year's LG Gram Pro lineup when they company unveiled the details last week. From the announcement, we found out that four new models are available. Here at CES 2025, I was able to check out the LG Gram Pro 2-in-1 in person to see what I was able to learn beyond the press release.