Goto

Collaborating Authors

 Large Language Model


Can Your Model Tell a Negation from an Implicature? Unravelling Challenges With Intent Encoders

arXiv.org Artificial Intelligence

Conversational systems often rely on embedding models for intent classification and intent clustering tasks. The advent of Large Language Models (LLMs), which enable instructional embeddings allowing one to adjust semantics over the embedding space using prompts, are being viewed as a panacea for these downstream conversational tasks. However, traditional evaluation benchmarks rely solely on task metrics that don't particularly measure gaps related to semantic understanding. Thus, we propose an intent semantic toolkit that gives a more holistic view of intent embedding models by considering three tasks -- (1) intent classification, (2) intent clustering, and (3) a novel triplet task. The triplet task gauges the model's understanding of two semantic concepts paramount in real-world conversational systems -- negation and implicature. We observe that current embedding models fare poorly in semantic understanding of these concepts. To address this, we propose a pre-training approach to improve the embedding model by leveraging augmentation with data generated by an auto-regressive model and a contrastive loss term. Our approach improves the semantic understanding of the intent embedding model on the aforementioned linguistic dimensions while slightly effecting their performance on downstream task metrics.


Modeling User Viewing Flow Using Large Language Models for Article Recommendation

arXiv.org Artificial Intelligence

This paper proposes the User Viewing Flow Modeling (SINGLE) method for the article recommendation task, which models the user constant preference and instant interest from user-clicked articles. Specifically, we first employ a user constant viewing flow modeling method to summarize the user's general interest to recommend articles. In this case, we utilize Large Language Models (LLMs) to capture constant user preferences from previously clicked articles, such as skills and positions. Then we design the user instant viewing flow modeling method to build interactions between user-clicked article history and candidate articles. It attentively reads the representations of user-clicked articles and aims to learn the user's different interest views to match the candidate article. Our experimental results on the Alibaba Technology Association (ATA) website show the advantage of SINGLE, achieving a 2.4% improvement over previous baseline models in the online A/B test. Our further analyses illustrate that SINGLE has the ability to build a more tailored recommendation system by mimicking different article viewing behaviors of users and recommending more appropriate and diverse articles to match user interests.


A Multimodal In-Context Tuning Approach for E-Commerce Product Description Generation

arXiv.org Artificial Intelligence

In this paper, we propose a new setting for generating product descriptions from images, augmented by marketing keywords. It leverages the combined power of visual and textual information to create descriptions that are more tailored to the unique features of products. For this setting, previous methods utilize visual and textual encoders to encode the image and keywords and employ a language model-based decoder to generate the product description. However, the generated description is often inaccurate and generic since same-category products have similar copy-writings, and optimizing the overall framework on large-scale samples makes models concentrate on common words yet ignore the product features. To alleviate the issue, we present a simple and effective Multimodal In-Context Tuning approach, named ModICT, which introduces a similar product sample as the reference and utilizes the in-context learning capability of language models to produce the description. During training, we keep the visual encoder and language model frozen, focusing on optimizing the modules responsible for creating multimodal in-context references and dynamic prompts. This approach preserves the language generation prowess of large language models (LLMs), facilitating a substantial increase in description diversity. To assess the effectiveness of ModICT across various language model scales and types, we collect data from three distinct product categories within the E-commerce domain. Extensive experiments demonstrate that ModICT significantly improves the accuracy (by up to 3.3% on Rouge-L) and diversity (by up to 9.4% on D-5) of generated results compared to conventional methods. Our findings underscore the potential of ModICT as a valuable tool for enhancing automatic generation of product descriptions in a wide range of applications. Code is at: https://github.com/HITsz-TMG/Multimodal-In-Context-Tuning


On Trojan Signatures in Large Language Models of Code

arXiv.org Artificial Intelligence

Trojan signatures, as described by Fields et al. (2021), are noticeable differences in the distribution of the trojaned class parameters (weights) and the non-trojaned class parameters of the trojaned model, that can be used to detect the trojaned model. Fields et al. (2021) found trojan signatures in computer vision classification tasks with image models, such as, Resnet, WideResnet, Densenet, and VGG. In this paper, we investigate such signatures in the classifier layer parameters of large language models of source code. Our results suggest that trojan signatures could not generalize to LLMs of code. We found that trojaned code models are stubborn, even when the models were poisoned under more explicit settings (finetuned with pre-trained weights frozen). We analyzed nine trojaned models for two binary classification tasks: clone and defect detection. To the best of our knowledge, this is the first work to examine weight-based trojan signature revelation techniques for large-language models of code and furthermore to demonstrate that detecting trojans only from the weights in such models is a hard problem.


Fine Tuning vs. Retrieval Augmented Generation for Less Popular Knowledge

arXiv.org Artificial Intelligence

Large language models (LLMs) memorize a vast amount of factual knowledge, exhibiting strong performance across diverse tasks and domains. However, it has been observed that the performance diminishes when dealing with less-popular or low-frequency concepts and entities, for example in domain specific applications. The two prominent approaches to enhance the performance of LLMs on low-frequent topics are: Retrieval Augmented Generation (RAG) and fine-tuning (FT) over synthetic data. This paper explores and evaluates the impact of RAG and FT on customizing LLMs in handling low-frequency entities on question answering task. Our findings indicate that FT significantly boosts the performance across entities of varying popularity, especially in the most and least popular groups, while RAG surpasses other methods. Additionally, the success of both RAG and FT approaches is amplified by advancements in retrieval and data augmentation techniques. We release our data and code at https://github.com/informagi/RAGvsFT.


The Lifeblood of the AI Boom

The Atlantic - Technology

Artificial intelligence can appear to be many different things--a whole host of programs with seemingly little common ground. Sometimes AI is a conversation partner, an illustrator, a math tutor, a facial-recognition tool. But in every incarnation, it is always, always a machine, demanding almost unfathomable amounts of data and energy to function. AI systems such as ChatGPT operate out of buildings stuffed with silicon computer chips. To build bigger machines--as Microsoft, Google, Meta, Amazon, and other tech companies would like to do--you need more resources.


Microsoft asks to dismiss New York Times's 'doomsday' copyright lawsuit

The Guardian

The tech giant said the lawsuit was near-sighted and akin to Hollywood's losing backlash against the VCR. In a motion to dismiss part of the lawsuit filed Monday, Microsoft, which was sued in December alongside ChatGPT-maker OpenAI, scoffed at the newspaper's claim that Times content receives "particular emphasis" and that tech companies "seek to free-ride on the Times's massive investment in its journalism". But in its response, Microsoft said the lawsuit was akin to Hollywood's resistance to the VCR that consumers used to record TV shows and which the entertainment business in the late 1970s feared would destroy its economic model. "'The VCR is to the American film producer and the American public as the Boston strangler is to the woman home alone,'" Microsoft said in its response, quoting from congressional testimony delivered by Jack Valenti, then head of the motion picture association of America, in 1982. In this case, Microsoft said, the Times is attempting to use "its might and its megaphone to challenge the latest profound technological advance: the Large Language Model."


Researchers Develop New Technique to Wipe Dangerous Knowledge From AI Systems

TIME - Tech

A study published Tuesday provides a newly-developed way to measure whether an AI model contains potentially hazardous knowledge, along with a technique for removing the knowledge from an AI system while leaving the rest of the model relatively intact. Together, the findings could help prevent AI models from being used to carry out cyberattacks and deploy bioweapons. The study was conducted by researchers from Scale AI, an AI training data provider, and the Center for AI Safety, a nonprofit, along with a consortium of more than 20 experts in biosecurity, chemical weapons, and cybersecurity. The subject matter experts generated a set of questions that, taken together, could assess whether an AI model can assist in efforts to create and deploy weapons of mass destruction. The researchers from the Center for AI Safety, building on previous work that helps to understand how AI models represent concepts, developed the "mind wipe" technique.


The Science of Detecting LLM-Generated Text

Communications of the ACM

Recent advancements in natural language generation (NLG) technology have significantly improved the diversity, control, and quality of large language models (LLM)-generated text. A notable example is OpenAI's ChatGPT, which demonstrates exceptional performance in tasks such as answering questions, composing email messages, essays, and codes. However, this newfound capability to produce human-like text at high efficiency also raises concerns about detecting and preventing misuse of LLMs in tasks such as phishing, disinformation, and academic dishonesty. For instance, many schools banned ChatGPT due to concerns over cheating in assignments,11 and media outlets have raised the alarm over fake news generated by LLMs.14 These concerns about the misuse of LLMs have hindered the NLG application in important domains such as media and education.


OpenAI Says Musk Agreed the ChatGPT Maker Should Become a For-Profit Company

TIME - Tech

Elon Musk supported making OpenAI a for-profit company, the ChatGPT maker said, attacking a lawsuit from the wealthy investor who has accused the artificial intelligence business of betraying its founding goal to benefit humanity as it pursued profits instead. In its first response since the Tesla CEO sued last week, OpenAI vowed to get the claim thrown out and released emails from Musk, escalating the feud between the San Francisco-based company and the billionaire that bankrolled its creation years ago. "The mission of OpenAI is to ensure AGI benefits all of humanity, which means both building safe and beneficial AGI and helping create broadly distributed benefits," OpenAI said in a blog post late Tuesday from five company executives and computer scientists, including CEO Sam Altman. "We intend to move to dismiss all of Elon's claims." AGI refers to artificial general intelligence, which are general purpose AI systems that can perform just as well as -- or even better than -- humans in a wide variety of tasks.