Media
Survey on Abstractive Text Summarization: Dataset, Models, and Metrics
Nnadi, Gospel Ozioma, Bertini, Flavio
Readers and scholars often desire a concise summary (Too Long; Didn't Read - TL;DR) of texts to effectively prioritize information. However, creating document summaries is mentally taxing and time-consuming, especially considering the overwhelming volume of documents produced annually, as depicted in Figure 1 by [2], Figure 2, [3] reported over 100,000 scientific articles on the Corona virus pandemic in 2020, though these articles contain brief abstracts of the article, the sheer volume poses challenges for researchers and medical professionals in quickly extracting relevant knowledge on a specific topic. An automatically generated multi-document summarization could be valuable, providing readers with essential information and reducing the need to access original files unless refinement is necessary. Text summarization has garnered significant research attention, proving useful in search engines, news clustering, timeline generation, and various other applications. The objective of text summarization is to create a brief, coherent, factually consistent, and readable document that retains the essential information from the source document, whether it is a single or multi-document. In Single Document Summarization (SDS) only one input document is used, eliminating the need for additional processing to assess relationships between inputs. This method is suitable for summarizing standalone documents such as emails, legal contracts, financial reports and so on. The primary goal of Multi Document Summarization (MDS) is to gather information from several texts addressing the same topic, often composed at different times or representing diverse perspectives. The overarching objective is to produce information reports that are both succinct and comprehensive, consolidating varied opinions from documents that explore a topic through multiple viewpoints.
Semantic Web: Past, Present, and Future
Scherp, Ansgar, Groener, Gerd, ล koda, Petr, Hose, Katja, Vidal, Maria-Esther
Ever since the vision was formulated, the Semantic Web has inspired many generations of innovations. Semantic technologies have been used to share vast amounts of information on the Web, enhance them with semantics to give them meaning, and enable inference and reasoning on them. Throughout the years, semantic technologies, and in particular knowledge graphs, have been used in search engines, data integration, enterprise settings, and machine learning. In this paper, we recap the classical concepts and foundations of the Semantic Web as well as modern and recent concepts and applications, building upon these foundations. The classical topics we cover include knowledge representation, creating and validating knowledge on the Web, reasoning and linking, and distributed querying. We enhance this classical view of the so-called ``Semantic Web Layer Cake'' with an update of recent concepts that include provenance, security and trust, as well as a discussion of practical impacts from industry-led contributions. We conclude with an outlook on the future directions of the Semantic Web.
A Reality Check on Context Utilisation for Retrieval-Augmented Generation
Hagstrรถm, Lovisa, Marjanoviฤ, Sara Vera, Yu, Haeun, Arora, Arnav, Lioma, Christina, Maistro, Maria, Atanasova, Pepa, Augenstein, Isabelle
Retrieval-augmented generation (RAG) helps address the limitations of the parametric knowledge embedded within a language model (LM). However, investigations of how LMs utilise retrieved information of varying complexity in real-world scenarios have been limited to synthetic contexts. We introduce DRUID (Dataset of Retrieved Unreliable, Insufficient and Difficult-to-understand contexts) with real-world queries and contexts manually annotated for stance. The dataset is based on the prototypical task of automated claim verification, for which automated retrieval of real-world evidence is crucial. We compare DRUID to synthetic datasets (CounterFact, ConflictQA) and find that artificial datasets often fail to represent the complex and diverse real-world context settings. We show that synthetic datasets exaggerate context characteristics rare in real retrieved data, which leads to inflated context utilisation results, as measured by our novel ACU score. Moreover, while previous work has mainly focused on singleton context characteristics to explain context utilisation, correlations between singleton context properties and ACU on DRUID are surprisingly small compared to other properties related to context source. Overall, our work underscores the need for real-world aligned context utilisation studies to represent and improve performance in real-world RAG settings.
Robustness of Large Language Models Against Adversarial Attacks
Tao, Yiyi, Shen, Yixian, Zhang, Hang, Shen, Yanxin, Wang, Lun, Shi, Chuanqi, Du, Shaoshuai
In this paper, we present a comprehensive study on the robustness of GPT LLM family. We employ two distinct evaluation methods to assess their resilience. The first method introduce character-level text attack in input prompts, testing the models on three sentiment classification datasets: StanfordNLP/IMDB, Yelp Reviews, and SST-2. The second method involves using jailbreak prompts to challenge the safety mechanisms of the LLMs. Our experiments reveal significant variations in the robustness of these models, demonstrating their varying degrees of vulnerability to both character-level and semantic-level adversarial attacks. These findings underscore the necessity for improved adversarial training and enhanced safety mechanisms to bolster the robustness of LLMs.
LLM-Powered User Simulator for Recommender System
Zhang, Zijian, Liu, Shuchang, Liu, Ziru, Zhong, Rui, Cai, Qingpeng, Zhao, Xiangyu, Zhang, Chunxu, Liu, Qidong, Jiang, Peng
User simulators can rapidly generate a large volume of timely user behavior data, providing a testing platform for reinforcement learning-based recommender systems, thus accelerating their iteration and optimization. However, prevalent user simulators generally suffer from significant limitations, including the opacity of user preference modeling and the incapability of evaluating simulation accuracy. In this paper, we introduce an LLM-powered user simulator to simulate user engagement with items in an explicit manner, thereby enhancing the efficiency and effectiveness of reinforcement learning-based recommender systems training. Specifically, we identify the explicit logic of user preferences, leverage LLMs to analyze item characteristics and distill user sentiments, and design a logical model to imitate real human engagement. By integrating a statistical model, we further enhance the reliability of the simulation, proposing an ensemble model that synergizes logical and statistical insights for user interaction simulations. Capitalizing on the extensive knowledge and semantic generation capabilities of LLMs, our user simulator faithfully emulates user behaviors and preferences, yielding high-fidelity training data that enrich the training of recommendation algorithms. We establish quantifying and qualifying experiments on five datasets to validate the simulator's effectiveness and stability across various recommendation scenarios.
Fox News AI Newsletter: Cate Blanchett 'deeply concerned'
LONDON, ENGLAND - DECEMBER 03: Cate Blanchett attends the World Premiere of "The Lord Of The Rings: The War Of The Rohirrim" at Odeon Luxe Leicester Square on December 3, 2024 in London, England. 'DEEPLY CONCERNED': Cate Blanchett is one of the many actors expressing fears about artificial intelligence. In a recent interview with the BBC, the Oscar winner said the technology "deeply concerned" her. ALTMAN OPENS UP: OpenAI CEO and co-founder Sam Altman opened up about Elon Musk's feud with him and his view of how regulations related to artificial intelligence development should be framed. CHATBOT SAFETY: This is a heartbreaking story out of Florida.
"Babygirl" Never Really Makes a Mess
In November, the reality star and entrepreneur Kim Kardashian posted a series of images and videos to her social-media accounts, in which she appeared to promote Tesla's new A.I. robot, Optimus. In a video on X, captioned "Meet my new friend," Kardashian is seen engaging with Elon Musk's humanoid golem, which reportedly retails for around thirty thousand dollars, and whose metal torso is inscribed with the Tesla logo. "O.K., hi!" she says perkily, off camera, as she waves her manicured fingers just within frame--a motion that is immediately echoed by the robot. "Can you do this: 'I love you'?" she asks next, forming a half heart with her hand, proffering it to the robot to urge him to complete the shape, and gasping in awe as he eagerly complies. But Optimus, who in the video seems more than happy to be at his mistress's beck and call, appears less subservient in a series of pictures in which Kardashian, wearing spike heels and lingerie, poses beside him and a gold Tesla Cybercab.
Music Can Thrive in the AI Era
The birth of ChatGPT brought a collection of anxieties regarding how large language models allow users to quickly subvert processes that once required human time, effort, passion, and understanding. And further, the tech sector's often stormy relationship with regulation and ethical oversight have left many fearful for a future where artificial intelligence replaces humans at work and stymies human creativity. While much of this alarm is well founded, we should also consider the possibility that human creativity can blossom in the age of AI. In 2025, we will start to see this manifest in our collective cultural response to technology. To examine how culture and creativity might adapt to the age of AI, we'll use hip-hop as an example.
DragonVerseQA: Open-Domain Long-Form Context-Aware Question-Answering
Lahiri, Aritra Kumar, Hu, Qinmin Vivian
This paper proposes a novel approach to develop an open-domain and long-form Over-The-Top (OTT) Question-Answering (QA) dataset, DragonVerseQA, specifically oriented to the fantasy universe of "House of the Dragon" and "Game Of Thrones" TV series. Most existing QA datasets focus on short, fact-based answers sourced almost solely from Wikipedia articles, devoid of depth and contextual richness for sophisticated narrative understanding. We curate a dataset that combines full episode summaries sourced from HBO and fandom wiki websites, user reviews from sources like IMDb and Rotten Tomatoes, and high-quality, open-domain, legally admissible sources, and structured data from repositories like WikiData into one dataset. The dataset provides a multi-dimensional context, reflecting complex character dynamics and plot developments from these varied sources. That means, on equal footing, only after heavy data preprocessing and filtering methods will meaningful, non-spam unbiased reviews be available in this enriched dataset. The comprehensive insights are given through the long-form answers generated from this enriched context. This is what makes this valuable dataset for improving conversational AI, narrative analysis, sentiment analysis, summarization techniques, and relation extraction. A comparative analysis with state-of-the-art QA datasets such as SQuAD 2.0, TriviaQA, and Natural Questions brings to light the unique advantages of our dataset in terms of contextual complexity and answer length. Detailed reviews add layers to audience sentiment and narrative interpretation, raising the bar for domain-specific QA with a new quality benchmark. Our work also allows a deeper understanding of entertainment-industry content and opens the door to more knowledgeable and creative AI-driven interactions within digital media environments.
Back To The Future: A Hybrid Transformer-XGBoost Model for Action-oriented Future-proofing Nowcasting
The interplay between past, present, and future is a central theme in the iconic movie Back to the Future, where small alterations in past events have profound, cascading effects on the future [1]. This concept mirrors the intricate and often non-linear relationships in real-world systems, where predictions about the future are not merely passive observations but active drivers of current decisions and behaviors. In the film, the characters reshape their present and future by altering past events, reflecting the power of temporal causality--the idea that events in time are interconnected, and that actions taken now have consequences for the future. In much the same way, effective nowcasting--predicting short-term outcomes like weather, natural hazards, health events, or traffic patterns--should not only anticipate what will happen but also incorporate how those predictions can influence present decisions and conditions. Traditional nowcasting methods, however, often focus exclusively on making predictions about future states without considering the active feedback loop that can be created by those predictions [2].