Goto

Collaborating Authors

 Media


Ad Auctions for LLMs via Retrieval Augmented Generation

arXiv.org Artificial Intelligence

The emergence of AI-driven assistant models like ChatGPT, Gemini, and Claude has influenced how individuals interact with these technologies, increasingly using them to streamline and enhance their work. While LLMs provide a fresh way to engage with information, the most advanced models are costly to operate (Minaee et al., 2024). To date, online advertising has been one of the most successful business models of the digital economy. Ads support a wide variety of online content and services, ranging from search engines, online publishers, to video content and more. However, LLM services today predominantly follow a subscription model (OpenAI, 2024). A natural question to ask in this context is whether advertising could support LLMs to alleviate serving costs and charges to users, and what format advertising on LLMs might take. In this paper, we develop auctions that allocate online ads within the output of LLMs using the framework of retrieval augmented generation (RAG) (Lewis et al., 2020). RAG is one of the most popular techniques to integrate factual information into the output of LLMs.


Siri gets overhaul as Apple goes all in on AI connected to ChatGPT

FOX News

Kurt Knutsson on the latest Apple updates that will include AI integration, prompting security concerns from Elon Musk. Apple held its annual developer's conference on Monday, announcing new software upgrades for all of its devices. IOS, which is the operating system that runs on your iPhone, has received what can be considered the biggest upgrade to date. Apple has infused it with artificial intelligence, meaning it is now more capable and feature-rich. IOS 18 is also more customizable than ever, giving you the ability to tweak your home screen and more.


Apple enters the AI race: How the tech giant is embedding artificial intelligence across ALL your devices and apps whether you like it or not - including removing people from your photos and tracking your family on flights

Daily Mail - Science & tech

After months of silence on its AI ambitions, Apple entered the artificial intelligence race with a lavish product announcement on Monday. At its Worldwide Developer Conference (WWDC), the multi-trillion dollar tech giant heralded a new era of technology, dubbed'Apple Intelligence'. Apple Intelligence is essentially a snazzy brand name for Apple's new-found focus on AI, triggered by the huge success of the ChatGPT chatbot 18 months ago. It means there will be an extensive presence of AI across Apple's devices and apps – whether you like it or not. While Apple claims the technology will usher in a'new chapter in Apple innovation', it seems that not everyone agrees, with Elon Musk dramatically warning that he will ban Apple devices from his firms following the news.


Danish Media Threatens to Sue OpenAI

WIRED

In the latest battle between AI and the media, major Danish newspapers and TV stations are threatening to sue OpenAI unless the company compensates the country's press for allegedly using their content to train its models. "We want remuneration for our work [which] they have used to train their model," says Karen Rønde, CEO of the Danish Press Publications' Collective Management Organization (DPCMO), which represents 99 percent of Danish media outlets, including state broadcaster DR and TV 2. Rønde says the DPCMO plans to sue if a deal is not reached in the next year. Soon after those lawsuits, OpenAI struck a series of licensing deals with major publishers, enabling the company to train its future iterations of ChatGPT on their content. Financial terms for the deals have not been disclosed. Now, Danish media is attempting to force OpenAI to negotiate with them as a collective, an unusual tactic that could provide a model for other small countries if successful.


Toxic Memes: A Survey of Computational Perspectives on the Detection and Explanation of Meme Toxicities

arXiv.org Artificial Intelligence

Internet memes, channels for humor, social commentary, and cultural expression, are increasingly used to spread toxic messages. Studies on the computational analyses of toxic memes have significantly grown over the past five years, and the only three surveys on computational toxic meme analysis cover only work published until 2022, leading to inconsistent terminology and unexplored trends. Our work fills this gap by surveying content-based computational perspectives on toxic memes, and reviewing key developments until early 2024. Employing the PRISMA methodology, we systematically extend the previously considered papers, achieving a threefold result. First, we survey 119 new papers, analyzing 158 computational works focused on content-based toxic meme analysis. We identify over 30 datasets used in toxic meme analysis and examine their labeling systems. Second, after observing the existence of unclear definitions of meme toxicity in computational works, we introduce a new taxonomy for categorizing meme toxicity types. We also note an expansion in computational tasks beyond the simple binary classification of memes as toxic or non-toxic, indicating a shift towards achieving a nuanced comprehension of toxicity. Third, we identify three content-based dimensions of meme toxicity under automatic study: target, intent, and conveyance tactics. We develop a framework illustrating the relationships between these dimensions and meme toxicities. The survey analyzes key challenges and recent trends, such as enhanced cross-modal reasoning, integrating expert and cultural knowledge, the demand for automatic toxicity explanations, and handling meme toxicity in low-resource languages. Also, it notes the rising use of Large Language Models (LLMs) and generative AI for detecting and generating toxic memes. Finally, it proposes pathways for advancing toxic meme detection and interpretation.


On the Robustness of Document-Level Relation Extraction Models to Entity Name Variations

arXiv.org Artificial Intelligence

Driven by the demand for cross-sentence and large-scale relation extraction, document-level relation extraction (DocRE) has attracted increasing research interest. Despite the continuous improvement in performance, we find that existing DocRE models which initially perform well may make more mistakes when merely changing the entity names in the document, hindering the generalization to novel entity names. To this end, we systematically investigate the robustness of DocRE models to entity name variations in this work. We first propose a principled pipeline to generate entity-renamed documents by replacing the original entity names with names from Wikidata. By applying the pipeline to DocRED and Re-DocRED datasets, we construct two novel benchmarks named Env-DocRED and Env-Re-DocRED for robustness evaluation. Experimental results show that both three representative DocRE models and two in-context learned large language models consistently lack sufficient robustness to entity name variations, particularly on cross-sentence relation instances and documents with more entities. Finally, we propose an entity variation robust training method which not only improves the robustness of DocRE models but also enhances their understanding and reasoning capabilities. We further verify that the basic idea of this method can be effectively transferred to in-context learning for DocRE as well.


Labeling Comic Mischief Content in Online Videos with a Multimodal Hierarchical-Cross-Attention Model

arXiv.org Artificial Intelligence

We address the challenge of detecting questionable content in online media, specifically the subcategory of comic mischief. This type of content combines elements such as violence, adult content, or sarcasm with humor, making it difficult to detect. Employing a multimodal approach is vital to capture the subtle details inherent in comic mischief content. To tackle this problem, we propose a novel end-to-end multimodal system for the task of comic mischief detection. As part of this contribution, we release a novel dataset for the targeted task consisting of three modalities: video, text (video captions and subtitles), and audio. We also design a HIerarchical Cross-attention model with CAPtions (HICCAP) to capture the intricate relationships among these modalities. The results show that the proposed approach makes a significant improvement over robust baselines and state-of-the-art models for comic mischief detection and its type classification. This emphasizes the potential of our system to empower users, to make informed decisions about the online content they choose to see. In addition, we conduct experiments on the UCF101, HMDB51, and XD-Violence datasets, comparing our model against other state-of-the-art approaches showcasing the outstanding performance of our proposed model in various scenarios.


BCAmirs at SemEval-2024 Task 4: Beyond Words: A Multimodal and Multilingual Exploration of Persuasion in Memes

arXiv.org Artificial Intelligence

Memes, combining text and images, frequently use metaphors to convey persuasive messages, shaping public opinion. Motivated by this, our team engaged in SemEval-2024 Task 4, a hierarchical multi-label classification task designed to identify rhetorical and psychological persuasion techniques embedded within memes. To tackle this problem, we introduced a caption generation step to assess the modality gap and the impact of additional semantic information from images, which improved our result. Our best model utilizes GPT-4 generated captions alongside meme text to fine-tune RoBERTa as the text encoder and CLIP as the image encoder. It outperforms the baseline by a large margin in all 12 subtasks. In particular, it ranked in top-3 across all languages in Subtask 2a, and top-4 in Subtask 2b, demonstrating quantitatively strong performance. The improvement achieved by the introduced intermediate step is likely attributable to the metaphorical essence of images that challenges visual encoders. This highlights the potential for improving abstract visual semantics encoding.


Interactive Perception for Deformable Object Manipulation

arXiv.org Artificial Intelligence

Interactive perception enables robots to manipulate the environment and objects to bring them into states that benefit the perception process. Deformable objects pose challenges to this due to significant manipulation difficulty and occlusion in vision-based perception. In this work, we address such a problem with a setup involving both an active camera and an object manipulator. Our approach is based on a sequential decision-making framework and explicitly considers the motion regularity and structure in coupling the camera and manipulator. We contribute a method for constructing and computing a subspace, called Dynamic Active Vision Space (DAVS), for effectively utilizing the regularity in motion exploration. The effectiveness of the framework and approach are validated in both a simulation and a real dual-arm robot setup. Our results confirm the necessity of an active camera and coordinative motion in interactive perception for deformable objects.


Learning Domain-Invariant Features for Out-of-Context News Detection

arXiv.org Artificial Intelligence

Multimodal out-of-context news is a common type of misinformation on online media platforms. This involves posting a caption, alongside an invalid out-of-context news image. Reflecting its importance, researchers have developed models to detect such misinformation. However, a common limitation of these models is that they only consider the scenario where pre-labeled data is available for each domain, failing to address the out-of-context news detection on unlabeled domains (e.g., unverified news on new topics or agencies). In this work, we therefore focus on domain adaptive out-of-context news detection. In order to effectively adapt the detection model to unlabeled news topics or agencies, we propose ConDA-TTA (Contrastive Domain Adaptation with Test-Time Adaptation) which applies contrastive learning and maximum mean discrepancy (MMD) to learn the domain-invariant feature. In addition, it leverages target domain statistics during test-time to further assist domain adaptation. Experimental results show that our approach outperforms baselines in 5 out of 7 domain adaptation settings on two public datasets, by as much as 2.93% in F1 and 2.08% in accuracy.