Media
SWING: Balancing Coverage and Faithfulness for Dialogue Summarization
Huang, Kung-Hsiang, Singh, Siffi, Ma, Xiaofei, Xiao, Wei, Nan, Feng, Dingwall, Nicholas, Wang, William Yang, McKeown, Kathleen
Missing information is a common issue of dialogue summarization where some information in the reference summaries is not covered in the generated summaries. To address this issue, we propose to utilize natural language inference (NLI) models to improve coverage while avoiding introducing factual inconsistencies. Specifically, we use NLI to compute fine-grained training signals to encourage the model to generate content in the reference summaries that have not been covered, as well as to distinguish between factually consistent and inconsistent generated sentences. Experiments on the DialogSum and SAMSum datasets confirm the effectiveness of the proposed approach in balancing coverage and faithfulness, validated with automatic metrics and human evaluations. Additionally, we compute the correlation between commonly used automatic metrics with human judgments in terms of three different dimensions regarding coverage and factual consistency to provide insight into the most suitable metric for evaluating dialogue summaries.
BRAIN L: A book recommender system
Sujo, Jessie Caridad Martรญn, Ribรฉ, Elisabet Golobardes i
Book sales in Spain have fallen progressively, which requires urgent changes to optimize the sales process as much as possible. This research proposes a new system, called Base of Reasoning in Artificial Intelligence with Natural Language (BRAIN L) focused exclusively on the publishing industry. The new field of knowledge of Artificial Intelligence (AI), Natural Language Processing (NLP), tecnolog\'ia del Machine Learning is combined with Case-Based Reasoning (CBR) techniques for book recommendations. A model is developed to retrieve similar cases/books supported by NLP techniques for decision making. In addition, policies are implemented to keep the model evaluated by expert reviews, where the system not only learns with new cases, but these cases are real.
Generate rather than Retrieve: Large Language Models are Strong Context Generators
Yu, Wenhao, Iter, Dan, Wang, Shuohang, Xu, Yichong, Ju, Mingxuan, Sanyal, Soumya, Zhu, Chenguang, Zeng, Michael, Jiang, Meng
Knowledge-intensive tasks, such as open-domain question answering (QA), require access to a large amount of world or domain knowledge. A common approach for knowledge-intensive tasks is to employ a retrieve-then-read pipeline that first retrieves a handful of relevant contextual documents from an external corpus such as Wikipedia and then predicts an answer conditioned on the retrieved documents. In this paper, we present a novel perspective for solving knowledge-intensive tasks by replacing document retrievers with large language model generators. We call our method generate-then-read (GenRead), which first prompts a large language model to generate contextutal documents based on a given question, and then reads the generated documents to produce the final answer. Furthermore, we propose a novel clustering-based prompting method that selects distinct prompts, resulting in the generated documents that cover different perspectives, leading to better recall over acceptable answers. We conduct extensive experiments on three different knowledge-intensive tasks, including open-domain QA, fact checking, and dialogue system. Notably, GenRead achieves 71.6 and 54.4 exact match scores on TriviaQA and WebQ, significantly outperforming the state-of-the-art retrieve-then-read pipeline DPR-FiD by +4.0 and +3.9, without retrieving any documents from any external knowledge source. Lastly, we demonstrate the model performance can be further improved by combining retrieval and generation. Our code and generated documents can be found at https://github.com/wyu97/GenRead.
Conversational Information Seeking
Zamani, Hamed, Trippas, Johanne R., Dalton, Jeff, Radlinski, Filip
Conversational information seeking (CIS) is concerned with a sequence of interactions between one or more users and an information system. Interactions in CIS are primarily based on natural language dialogue, while they may include other types of interactions, such as click, touch, and body gestures. This monograph provides a thorough overview of CIS definitions, applications, interactions, interfaces, design, implementation, and evaluation. This monograph views CIS applications as including conversational search, conversational question answering, and conversational recommendation. Our aim is to provide an overview of past research related to CIS, introduce the current state-of-the-art in CIS, highlight the challenges still being faced in the community. and suggest future directions.
VAuLT: Augmenting the Vision-and-Language Transformer for Sentiment Classification on Social Media
Chochlakis, Georgios, Srinivasan, Tejas, Thomason, Jesse, Narayanan, Shrikanth
We propose the Vision-and-Augmented-Language Transformer (VAuLT). VAuLT is an extension of the popular Vision-and-Language Transformer (ViLT), and improves performance on vision-and-language (VL) tasks that involve more complex text inputs than image captions while having minimal impact on training and inference efficiency. ViLT, importantly, enables efficient training and inference in VL tasks, achieved by encoding images using a linear projection of patches instead of an object detector. However, it is pretrained on captioning datasets, where the language input is simple, literal, and descriptive, therefore lacking linguistic diversity. So, when working with multimedia data in the wild, such as multimodal social media data, there is a notable shift from captioning language data, as well as diversity of tasks. We indeed find evidence that the language capacity of ViLT is lacking. The key insight and novelty of VAuLT is to propagate the output representations of a large language model (LM) like BERT to the language input of ViLT. We show that joint training of the LM and ViLT can yield relative improvements up to 20% over ViLT and achieve state-of-the-art or comparable performance on VL tasks involving richer language inputs and affective constructs, such as for Target-Oriented Sentiment Classification in TWITTER-2015 and TWITTER-2017, and Sentiment Classification in MVSA-Single and MVSA-Multiple. Our code is available at https://github.com/gchochla/VAuLT.
Automated multilingual detection of Pro-Kremlin propaganda in newspapers and Telegram posts
Solopova, Veronika, Popescu, Oana-Iuliana, Benzmรผller, Christoph, Landgraf, Tim
The full-scale conflict between the Russian Federation and Ukraine generated an unprecedented amount of news articles and social media data reflecting opposing ideologies and narratives. These polarized campaigns have led to mutual accusations of misinformation and fake news, shaping an atmosphere of confusion and mistrust for readers worldwide. This study analyses how the media affected and mirrored public opinion during the first month of the war using news articles and Telegram news channels in Ukrainian, Russian, Romanian and English. We propose and compare two methods of multilingual automated pro-Kremlin propaganda identification, based on Transformers and linguistic features. We analyse the advantages and disadvantages of both methods, their adaptability to new genres and languages, and ethical considerations of their usage for content moderation. With this work, we aim to lay the foundation for further development of moderation tools tailored to the current conflict.
A Holistic Cascade System, benchmark, and Human Evaluation Protocol for Expressive Speech-to-Speech Translation
Huang, Wen-Chin, Peloquin, Benjamin, Kao, Justine, Wang, Changhan, Gong, Hongyu, Salesky, Elizabeth, Adi, Yossi, Lee, Ann, Chen, Peng-Jen
Expressive speech-to-speech translation (S2ST) aims to transfer prosodic attributes of source speech to target speech while maintaining translation accuracy. Existing research in expressive S2ST is limited, typically focusing on a single expressivity aspect at a time. Likewise, this research area lacks standard evaluation protocols and well-curated benchmark datasets. In this work, we propose a holistic cascade system for expressive S2ST, combining multiple prosody transfer techniques previously considered only in isolation. We curate a benchmark expressivity test set in the TV series domain and explored a second dataset in the audiobook domain. Finally, we present a human evaluation protocol to assess multiple expressive dimensions across speech pairs. Experimental results indicate that bi-lingual annotators can assess the quality of expressive preservation in S2ST systems, and the holistic modeling approach outperforms single-aspect systems. Audio samples can be accessed through our demo webpage: https://facebookresearch.github.io/speech_translation/cascade_expressive_s2st.
'Jung_E' Netflix Review: Stream It or Skip It?
Jung_E (now on Netflix) is the new film from director Yeon Sang-ho, who made a name for himself outside his native Korea with 2016 zombie action movie Train to Busan. As he offered a new angle on a familiar subgenre with Busan, he surely hopes to do the same for artificial intelligence science-fiction (AI-SCI-FI?!?) with his latest work, which is set in a sort-of-post-apocalyptic dystopia where robots fight wars for us, and the side with the best AI sure seems ripe for victory. The movie is also notable for being the final role of Korean film star Kang Soo-yeon, who sadly passed away in 2022 at age 55 after suffering a cerebral hemorrhage. The Gist: IN A WORLD where severe climate change has forced humanity to mostly abandon Earth for the Moon; where subsequent civil war has raged for decades; where artificial intelligence is a primary component of war technology; where human brains can be uploaded from diseased bodies to new ones and if you have enough money you can enjoy a terrific Type A existence, or a so-so Type B, or possibly a horrific, but no-cost Type C where your consciousness is under the control of corporations and shit; where people take ethics tests to determine that they're indeed human and not AI โ in this world, a woman leaps around a bona-fide Dystopian Hell of a set piece, fighting robots, some more diabolically advanced in their ability to withstand bullets and such. She is the famed kickass warrior Yun Jung-yi (Kim Hyun-joo), but she really isn't Yun Jung-yi โ she's Jung_E, a clone of Yun Jung-yi, and she keeps failing the same battle simulation. The simulation reinvents the very scene of her defeat many years prior, which left her body in a coma, the contents of her brain as the key element of weapons-development research firm Kronoid and her daughter kind of almost orphaned.
The 10 Best Shows on Apple TV Right Now
Slowly but surely Apple TV is finding its feet. The streaming service, which at launch we called "odd, angsty, and horny as hell," has evolved into a diverse library of dramas, documentaries, and comedies. It's also fairly cheap compared to services like Netflix--and Apple often throws in three free months when you buy a new iPhone, iPad, Mac, or Apple TV. Curious but don't know where to get started? Below are our picks for the best shows on the service.
1923 cartoon eerily predicted 2023's AI art generators
In 1923, an editorial cartoonist named H.T. Webster drew a humorous cartoon for the New York World newspaper depicting a fictional 2023 machine that would generate ideas and draw them as cartoons automatically. It presaged recent advancements in AI image synthesis, one century later, that actually can create artwork automatically. The vintage cartoon carries the caption "In the year 2023 when all our work is done by electricity." It depicts a cartoonist standing by his drawing table and making plans for social events while an "idea dynamo" generates ideas and a "cartoon dynamo" renders the artwork. Interestingly, this separation of labor feels similar to our neural networks of today.