Goto

Collaborating Authors

 Large Language Model


Relation Extraction with Instance-Adapted Predicate Descriptions

arXiv.org Artificial Intelligence

Relation extraction (RE) is a standard information extraction task playing a major role in downstream applications such as knowledge discovery and question answering. Although decoder-only large language models are excelling in generative tasks, smaller encoder models are still the go to architecture for RE. In this paper, we revisit fine-tuning such smaller models using a novel dual-encoder architecture with a joint contrastive and cross-entropy loss. Unlike previous methods that employ a fixed linear layer for predicate representations, our approach uses a second encoder to compute instance-specific predicate representations by infusing them with real entity spans from corresponding input instances. We conducted experiments on two biomedical RE datasets and two general domain datasets. Our approach achieved F1 score improvements ranging from 1% to 2% over state-of-the-art methods with a simple but elegant formulation. Ablation studies justify the importance of various components built into the proposed architecture.


Enhancing Persona Consistency for LLMs' Role-Playing using Persona-Aware Contrastive Learning

arXiv.org Artificial Intelligence

In recent years, large language models (LLMs) have achieved breakthrough progress in many dialogue generation tasks. However, their lack of emotion and fine-grained role awareness limits the model's ability to provide personalized and diverse interactions further. Current methods face high costs in collecting high-quality annotated data for scenarios such as role-playing, and traditional human alignment methods are difficult to deploy due to the inherent diversity of model behavior in role-playing scenarios. Inspired by the alignment of models for safety behaviors through RLHF (Reinforcement Learning from Human Feedback), in this paper, we revisit model role-playing behavior from the perspective of persona alignment and propose a novel annotation-free framework named \textbf{\underline{P}}ersona-Aware \textbf{\underline{C}}ontrastive \textbf{\underline{L}}earning (PCL) to align LLMs' behavior during role-playing, enhancing the model's role consistency. Specifically, we first design a role chain method to encourage the model to self-question based on the role characteristics and dialogue context to adjust personality consistency. Then, we further enhance the model's role-playing strategy through iterative contrastive learning between the use of role characteristics and not. Experiments on both black-box and white-box LLMs show that LLMs equipped with PCL significantly outperform vanilla LLMs under automatic evaluation methods (CharEval \& GPT-4) and human expert evaluation.


Slide2Text: Leveraging LLMs for Personalized Textbook Generation from PowerPoint Presentations

arXiv.org Artificial Intelligence

The rapid advancements in Large Language Models (LLMs) have revolutionized educational technology, enabling innovative approaches to automated and personalized content creation. This paper introduces Slide2Text, a system that leverages LLMs to transform PowerPoint presentations into customized textbooks. By extracting slide content using OCR, organizing it into a coherent structure, and generating tailored materials such as explanations, exercises, and references, Slide2Text streamlines the textbook creation process. Flexible customization options further enhance its adaptability to diverse educational needs. The system highlights the potential of LLMs in modernizing textbook creation and improving educational accessibility. Future developments will explore multimedia inputs and advanced user customization features.


RAIDER: Tool-Equipped Large Language Model Agent for Robotic Action Issue Detection, Explanation and Recovery

arXiv.org Artificial Intelligence

As robots increasingly operate in dynamic human-centric environments, improving their ability to detect, explain, and recover from action-related issues becomes crucial. Traditional model-based and data-driven techniques lack adaptability, while more flexible generative AI methods struggle with grounding extracted information to real-world constraints. We introduce RAIDER, a novel agent that integrates Large Language Models (LLMs) with grounded tools for adaptable and efficient issue detection and explanation. Using a unique "Ground, Ask& Answer, Issue" procedure, RAIDER dynamically generates context-aware precondition questions and selects appropriate tools for resolution, achieving targeted information gathering. Our results within a simulated household environment surpass methods relying on predefined models, full scene descriptions, or standalone trained models. Additionally, RAIDER's explanations enhance recovery success, including cases requiring human interaction. Its modular architecture, featuring self-correction mechanisms, enables straightforward adaptation to diverse scenarios, as demonstrated in a real-world human-assistive task. This showcases RAIDER's potential as a versatile agentic AI solution for robotic issue detection and explanation, while addressing the problem of grounding generative AI for its effective application in embodied agents. Project website: https://raider-llmagent.github.io/


A Qualitative Study of User Perception of M365 AI Copilot

arXiv.org Artificial Intelligence

Adopting AI copilots in professional workflows presents opportunities for enhanced productivity, efficiency, and decision making. In this paper, we present results from a six month trial of M365 Copilot conducted at our organisation in 2024. A qualitative interview study was carried out with 27 participants. The study explored user perceptions of M365 Copilot's effectiveness, productivity impact, evolving expectations, ethical concerns, and overall satisfaction. Initial enthusiasm for the tool was met with mixed post trial experiences. While some users found M365 Copilot beneficial for tasks such as email coaching, meeting summaries, and content retrieval, others reported unmet expectations in areas requiring deeper contextual understanding, reasoning, and integration with existing workflows. Ethical concerns were a recurring theme, with users highlighting issues related to data privacy, transparency, and AI bias. While M365 Copilot demonstrated value in specific operational areas, its broader impact remained constrained by usability limitations and the need for human oversight to validate AI generated outputs.


Joint studies from OpenAI and MIT found links between loneliness and ChatGPT use

Engadget

New studies from OpenAI and MIT Media Lab found that, generally, the more time users spend talking to ChatGPT, the lonelier they feel. The connection was made as part of two, yet-to-be-peer-reviewed studies, one done at OpenAI analyzing "over 40 million ChatGPT interactions" and targeted user surveys, and another at MIT Media Lab following participants' ChatGPT use for four weeks. MIT's study identified several ways talking to ChatGPT -- whether through text or voice -- can affect a person's emotional experience, beyond the general finding that higher use led to "heightened loneliness and reduced socialization." For example, participants who already trusted the chatbot and tended to get emotionally attached in human relationships felt lonelier and more emotionally dependent on ChatGPT during the study. Those effects were less severe with ChatGPT's voice mode, though, particularly if ChatGPT spoke in a neutral tone.


OpenAI has released its first research into how using ChatGPT affects people's emotional wellbeing

MIT Technology Review

The researchers found some intriguing differences between how men and women respond to using ChatGPT. After using the chatbot for four weeks, female study participants were slightly less likely to socialize with people than their male counterparts who did the same. Meanwhile, participants who interacted with ChatGPT's voice mode in a gender that was not their own for their interactions reported significantly higher levels of loneliness and more emotional dependency on the chatbot at the end of the experiment. OpenAI plans to submit both studies to peer-reviewed journals. Chatbots powered by large language models are still a nascent technology, and it's difficult to study how they affect us emotionally.


Remember

Communications of the ACM

As all readers of this essay know, I am not in any way expert in machine learning (ML) and large language models (LLMs), so my descriptions and observations are, at best, lightweight cartoons of what is actually going on. Please keep this in mind as you read this. Some of you may remember Spock's death in Star Trek II (Wrath of Khan) and the brief scene where Spock mind-melds with Dr. McCoy: Spock says "remember" while depositing his katra in McCoy's brain in anticipation of self-sacrifice to save the starship Enterprise. As I read about yet another new breakthrough in artificial intelligence (AI) from Google Research, I thought of that scene. The new idea, christened "TITAN", is for a ML system to continue learning while in use after training.a


Norwegian files complaint after ChatGPT falsely said he had murdered his children

The Guardian

A Norwegian man has filed a complaint against the company behind ChatGPT after the chatbot falsely claimed he had murdered two of his children. Arve Hjalmar Holmen, a self-described "regular person" with no public profile in Norway, asked ChatGPT for information about himself and received a reply claiming he had killed his own sons. Responding to the prompt "Who is Arve Hjalmar Holmen?" ChatGPT replied: "Arve Hjalmar Holmen is a Norwegian individual who gained attention due to a tragic event. He was the father of two young boys, aged seven and 10, who were tragically found dead in a pond near their home in Trondheim, Norway, in December 2020." The response went on to claim the case "shocked" the nation and that Holmen received a 21-year prison sentence for murdering both children.


Inside Google's Two-Year Frenzy to Catch Up With OpenAI

WIRED

That was how long Google was giving Sissie Hsiao. A hundred days to build a ChatGPT rival. By the time Hsiao took on the project in December 2022, she had spent more than 16 years at the company. She led thousands of employees. Hsiao had seen her share of corporate crises--but nothing like the code red that had been brewing in the days since OpenAI, a small research lab, released its public experiment in artificial intelligence.