Media
Online SuBmodular + SuPermodular (BP) Maximization with Bandit Feedback
Narang, Adhyyan, Sadeghi, Omid, Ratliff, Lillian J, Fazel, Maryam, Bilmes, Jeff
We investigate non-modular function maximization in an online setting with $m$ users. The optimizer maintains a set $S_q$ for each user $q \in \{1, \ldots, m\}$. At round $i$, a user with unknown utility $h_q$ arrives; the optimizer selects a new item to add to $S_q$, and receives a noisy marginal gain. The goal is to minimize regret compared to an $\alpha$-approximation to the optimal full-knowledge selection (i.e., $\alpha$-regret). Prior works study this problem under a submodularity assumption for all $h_q$. However, this is not ideally amenable to applications, e.g., movie recommendations, that involve complementarity between items, where e.g., watching the first movie in a series enhances the impression of watching the sequels. Hence, we consider objectives $h_q$, called \textit{BP functions}, that decompose into the sum of monotone submodular $f_q$ and supermodular $g_q$; here, $g_q$ naturally models complementarity. Under different feedback assumptions, we develop UCB-style algorithms that use Nystrom sampling for computational efficiency. For these, we provide sublinear $\alpha$-regret guarantees for $\alpha = 1/\kappa_{f} [1 - e^{-(1 - \kappa^g) \kappa_{f}} ]$, and $\alpha = \min\{1 - \kappa_f/e, 1 - \kappa^g\}$; here, $\kappa_f, \kappa^g$ are submodular and supermodular curvatures. Furthermore, we provide similar $\alpha$-regret guarantees for functions that are almost submodular where $\alpha$ is parameterized by the submodularity ratio of the objective functions. We numerically validate our algorithms for movie recommendation on the MovieLens dataset and selection of training subsets for classification tasks.
On Narrative Information and the Distillation of Stories
Ashley, Dylan R., Herrmann, Vincent, Friggstad, Zachary, Schmidhuber, Jรผrgen
The act of telling stories is a fundamental part of what it means to be human. This work introduces the concept of narrative information, which we define to be the overlap in information space between a story and the items that compose the story. Using contrastive learning methods, we show how modern artificial neural networks can be leveraged to distill stories and extract a representation of the narrative information. We then demonstrate how evolutionary algorithms can leverage this to extract a set of narrative templates and how these templates -- in tandem with a novel curve-fitting algorithm we introduce -- can reorder music albums to automatically induce stories in them. In the process of doing so, we give strong statistical evidence that these narrative information templates are present in existing albums. While we experiment only with music albums here, the premises of our work extend to any form of (largely) independent media.
Voice actors warn artificial intelligence could replace them, cut industry jobs and pay
Fox News host Steve Hilton delves into ChatGPT, an artificial intelligence program that could have major implications for writing-focused jobs on'The Next Revolution.' Actors are sounding the alarm on new artificial intelligence (AI) technology that creates replicas of their voices and could replace them without proper compensation. According to a report by VICE's Motherboard, Hollywood and videogame voice actors are being asked to sign contracts that give away the rights to their voices for use in generative AI. They claim that the increasingly common practice could decimate entire aspects of the industry. "It's disrespectful to the craft to suggest that generating a performance is equivalent to a real human being's performance," SungWon Cho, a game and animation voice actor, told Motherboard.
Eighteen pitfalls to beware of in AI journalism
Reporting about AI is hard. When news articles uncritically repeat PR statements, overuse images of robots, attribute agency to AI tools, or downplay their limitations, they mislead and misinform readers about the potential and limitations of AI. We noticed that many articles tend to mislead in similar ways, so we analyzed over 50 articles about AI from major publications, from which we compiled 18 recurring pitfalls. We hope that being familiar with these will help you detect hype whenever you see it. We also hope this compilation of pitfalls will help journalists avoid them.
Artificial intelligence makes voice cloning easy and 'the monster is already on the loose'
Digital forensics experts say the video was created using a new generation of artificial intelligence tools, which allow anyone to quickly generate audio simulating a person's voice with a few clicks of a button. And while the Biden clip on social media may have failed to fool most users this time, the clip shows how easy it now is for people to generate hateful and disinformation-filled "deepfake" videos that could do real-world harm. "Tools like this are going to basically add more fuel to fire," said Hafiz Malik, a professor of electrical and computer engineering at the University of Michigan who focuses on multimedia forensics. "The monster is already on the loose." It arrived last month with the beta phase of ElevenLabs' voice synthesis platform, which allowed users to generate realistic audio of any person's voice by uploading a few minutes of audio samples and typing in any text for it to say.
What's The Difference Between Machine Learning And Artificial Intelligence, Anyway?
If you pick up your phone and open a news app today, you're likely to come across some mention of artificial intelligence (AI). While the team at Q.ai has been working hard at using AI to manage investments for years, new developments like ChatGPT and Dall-E are captivating computer users of all backgrounds. If you don't know exactly what artificial intelligence means and how it differs from the related machine learning (ML) technology, here's a closer look at what you need to know about them when searching for profitable investments. The phrase artificial intelligence likely brings up images of sci-fi movies where space-ship-controlling computers or robot maids turn violent and try to take over the world. The reality of AI is much more boring than an army of computerized robots, but it's an exciting time for new AI technologies.
ChatGPT has only been around for two months and is causing untold chaos
It's safe to say ChatGPT is causing chaos. The AI chatbot from OpenAI has only been around for two months and has already amassed more than one million users. Launched on November 30, the chatbot has impressed -- and riled -- many different people. Chatter about the new tech has stretched far beyond the business world and even managed to provoke the disdain of award-winning songwriter Nick Cave. ChatGPT has already been compared with the launch of the iPhone and the crypto boom but while the tech's long-term influence remains to be seen, people are already finding creative ways to use it.
Analyzing the Effectiveness of the Underlying Reasoning Tasks in Multi-hop Question Answering
Ho, Xanh, Nguyen, Anh-Khoa Duong, Sugawara, Saku, Aizawa, Akiko
To explain the predicted answers and evaluate the reasoning abilities of models, several studies have utilized underlying reasoning (UR) tasks in multi-hop question answering (QA) datasets. However, it remains an open question as to how effective UR tasks are for the QA task when training models on both tasks in an end-to-end manner. In this study, we address this question by analyzing the effectiveness of UR tasks (including both sentence-level and entity-level tasks) in three aspects: (1) QA performance, (2) reasoning shortcuts, and (3) robustness. While the previous models have not been explicitly trained on an entity-level reasoning prediction task, we build a multi-task model that performs three tasks together: sentence-level supporting facts prediction, entity-level reasoning prediction, and answer prediction. Experimental results on 2WikiMultiHopQA and HotpotQA-small datasets reveal that (1) UR tasks can improve QA performance. Using four debiased datasets that are newly created, we demonstrate that (2) UR tasks are helpful in preventing reasoning shortcuts in the multi-hop QA task. However, we find that (3) UR tasks do not contribute to improving the robustness of the model on adversarial questions, such as sub-questions and inverted questions. We encourage future studies to investigate the effectiveness of entity-level reasoning in the form of natural language questions (e.g., sub-question forms).
AI vs. Human -- Differentiation Analysis of Scientific Content Generation
Ma, Yongqiang, Liu, Jiawei, Yi, Fan, Cheng, Qikai, Huang, Yong, Lu, Wei, Liu, Xiaozhong
Recent neural language models have taken a significant step forward in producing remarkably controllable, fluent, and grammatical text. Although studies have found that AI-generated text is not distinguishable from human-written text for crowd-sourcing workers, there still exist errors in AI-generated text which are even subtler and harder to spot. We primarily focus on the scenario in which scientific AI writing assistant is deeply involved. First, we construct a feature description framework to distinguish between AI-generated text and human-written text from syntax, semantics, and pragmatics based on the human evaluation. Then we utilize the features, i.e., writing style, coherence, consistency, and argument logistics, from the proposed framework to analyze two types of content. Finally, we adopt several publicly available methods to investigate the gap of between AI-generated scientific text and human-written scientific text by AI-generated scientific text detection models. The results suggest that while AI has the potential to generate scientific content that is as accurate as human-written content, there is still a gap in terms of depth and overall quality. The AI-generated scientific content is more likely to contain errors in factual issues. We find that there exists a "writing style" gap between AI-generated scientific text and human-written scientific text. Based on the analysis result, we summarize a series of model-agnostic and distribution-agnostic features for detection tasks in other domains. Findings in this paper contribute to guiding the optimization of AI models to produce high-quality content and addressing related ethical and security concerns.
Team Triple-Check at Factify 2: Parameter-Efficient Large Foundation Models with Feature Representations for Multi-Modal Fact Verification
Du, Wei-Wei, Wu, Hong-Wei, Wang, Wei-Yao, Peng, Wen-Chih
Multi-modal fact verification has become an important but challenging issue on social media due to the mismatch between the text and images in the misinformation of news content, which has been addressed by considering cross-modalities to identify the veracity of the news in recent years. In this paper, we propose the Pre-CoFactv2 framework with new parameter-efficient foundation models for modeling fine-grained text and input embeddings with lightening parameters, multi-modal multi-type fusion for not only capturing relations for the same and different modalities but also for different types (i.e., claim and document), and feature representations for explicitly providing metadata for each sample. In addition, we introduce a unified ensemble method to boost model performance by adjusting the importance of each trained model with not only the weights but also the powers. Extensive experiments show that Pre-CoFactv2 outperforms Pre-CoFact by a large margin and achieved new state-of-the-art results at the Factify challenge at AAAI 2023. We further illustrate model variations to verify the relative contributions of different components. Our team won the first prize (F1-score: 81.82%) and we made our code publicly available