Media
FKA twigs Creates Deepfake AI Version of Herself With a Special Use in Mind
British singer-songwriter FKA twigs, born Tahliah Debrett Barnett, testified before the U.S. Senate Judiciary Subcommittee on Intellectual Property on Tuesday about the dangers of artificial intelligence. She relayed that she was especially concerned as an artist whose music and performances are used by third parties to train artificial intelligence models. She said that the power of this technology has become especially apparent to her as she has attempted to build a deepfake version of herself. "In the past year, I have developed my own deepfake version of myself that is not only trained in my personality, but also can use my exact tone of voice to speak many languages," the singer said in her statement. "I will be engaging my'AI twigs' later this year to extend my reach and handle my online social media interactions, whilst I continue to focus on my art from the comfort and solace of my studio."
Microsoft and OpenAI sued yet again by Chicago Tribune and New York Daily News
A group of publications that include the Chicago Tribune, New York Daily News and the Orlando Sentinel are suing Microsoft and OpenAI, as reported by The Verge. Their products can regurgitate Times' articles verbatim and can "mimic its expressive style," the publication said, even though they didn't have a prior licensing agreement. In a motion seeking to dismiss key parts of the lawsuit, Microsoft accused the Times of doomsday futurology by claiming that generative AI can pose a threat to independent journalism. ACG's newspapers complain of the same thing, that the companies' chatbots are reproducing their articles word-for-word shortly after they're published without a prominent link back to the sources. They included several examples in their complaint.
Brazil's last Japanese-language newspaper innovates to stay in print
Diario Brasil Nippou, the last remaining Japanese-language newspaper in Brazil, is struggling to keep its presses rolling. The South American country is home to the largest Japanese community outside the East Asian nation, with some 2.7 million Nikkei Japanese immigrants and their descendants. Behind the difficulties facing the paper is a decline in the number of subscribers, partly reflecting the aging of immigrants from Japan.
Uncovering Agendas: A Novel French & English Dataset for Agenda Detection on Social Media
Katsios, Gregorios, Sa, Ning, Bhaumik, Ankita, Strzalkowski, Tomek
The behavior and decision making of groups or communities can be dramatically influenced by individuals pushing particular agendas, e.g., to promote or disparage a person or an activity, to call for action, etc.. In the examination of online influence campaigns, particularly those related to important political and social events, scholars often concentrate on identifying the sources responsible for setting and controlling the agenda (e.g., public media). In this article we present a methodology for detecting specific instances of agenda control through social media where annotated data is limited or non-existent. By using a modest corpus of Twitter messages centered on the 2022 French Presidential Elections, we carry out a comprehensive evaluation of various approaches and techniques that can be applied to this problem. Our findings demonstrate that by treating the task as a textual entailment problem, it is possible to overcome the requirement for a large annotated training dataset.
Addressing Topic Granularity and Hallucination in Large Language Models for Topic Modelling
Mu, Yida, Bai, Peizhen, Bontcheva, Kalina, Song, Xingyi
Large language models (LLMs) with their strong zero-shot topic extraction capabilities offer an alternative to probabilistic topic modelling and closed-set topic classification approaches. As zero-shot topic extractors, LLMs are expected to understand human instructions to generate relevant and non-hallucinated topics based on the given documents. However, LLM-based topic modelling approaches often face difficulties in generating topics with adherence to granularity as specified in human instructions, often resulting in many near-duplicate topics. Furthermore, methods for addressing hallucinated topics generated by LLMs have not yet been investigated. In this paper, we focus on addressing the issues of topic granularity and hallucinations for better LLM-based topic modelling. To this end, we introduce a novel approach that leverages Direct Preference Optimisation (DPO) to fine-tune open-source LLMs, such as Mistral-7B. Our approach does not rely on traditional human annotation to rank preferred answers but employs a reconstruction pipeline to modify raw topics generated by LLMs, thus enabling a fast and efficient training and inference framework. Comparative experiments show that our fine-tuning approach not only significantly improves the LLM's capability to produce more coherent, relevant, and precise topics, but also reduces the number of hallucinated topics.
Identifying Fairness Issues in Automatically Generated Testing Content
Stowe, Kevin, Longwill, Benny, Francis, Alyssa, Aoyama, Tatsuya, Ghosh, Debanjan, Somasundaran, Swapna
Natural language generation tools are powerful and effective for generating content. However, language models are known to display bias and fairness issues, making them impractical to deploy for many use cases. We here focus on how fairness issues impact automatically generated test content, which can have stringent requirements to ensure the test measures only what it was intended to measure. Specifically, we review test content generated for a large-scale standardized English proficiency test with the goal of identifying content that only pertains to a certain subset of the test population as well as content that has the potential to be upsetting or distracting to some test takers. Issues like these could inadvertently impact a test taker's score and thus should be avoided. This kind of content does not reflect the more commonly-acknowledged biases, making it challenging even for modern models that contain safeguards. We build a dataset of 601 generated texts annotated for fairness and explore a variety of methods for classification: fine-tuning, topic-based classification, and prompting, including few-shot and self-correcting prompts. We find that combining prompt self-correction and few-shot learning performs best, yielding an F1 score of 0.79 on our held-out test set, while much smaller BERT- and topic-based models have competitive performance on out-of-domain data.
QuanTemp: A real-world open-domain benchmark for fact-checking numerical claims
V, Venktesh, Anand, Abhijit, Anand, Avishek, Setty, Vinay
Automated fact checking has gained immense interest to tackle the growing misinformation in the digital era. Existing systems primarily focus on synthetic claims on Wikipedia, and noteworthy progress has also been made on real-world claims. In this work, we release QuanTemp, a diverse, multi-domain dataset focused exclusively on numerical claims, encompassing temporal, statistical and diverse aspects with fine-grained metadata and an evidence collection without leakage. This addresses the challenge of verifying real-world numerical claims, which are complex and often lack precise information, not addressed by existing works that mainly focus on synthetic claims. We evaluate and quantify the limitations of existing solutions for the task of verifying numerical claims. We also evaluate claim decomposition based methods, numerical understanding based models and our best baselines achieves a macro-F1 of 58.32. This demonstrates that QuanTemp serves as a challenging evaluation set for numerical claim verification.
A Hong Kong Sign Language Corpus Collected from Sign-interpreted TV News
Niu, Zhe, Zuo, Ronglai, Mak, Brian, Wei, Fangyun
This paper introduces TVB-HKSL-News, a new Hong Kong sign language (HKSL) dataset collected from a TV news program over a period of 7 months. The dataset is collected to enrich resources for HKSL and support research in large-vocabulary continuous sign language recognition (SLR) and translation (SLT). It consists of 16.07 hours of sign videos of two signers with a vocabulary of 6,515 glosses (for SLR) and 2,850 Chinese characters or 18K Chinese words (for SLT). One signer has 11.66 hours of sign videos and the other has 4.41 hours. One objective in building the dataset is to support the investigation of how well large-vocabulary continuous sign language recognition/translation can be done for a single signer given a (relatively) large amount of his/her training data, which could potentially lead to the development of new modeling methods. Besides, most parts of the data collection pipeline are automated with little human intervention; we believe that our collection method can be scaled up to collect more sign language data easily for SLT in the future for any sign languages if such sign-interpreted videos are available. We also run a SOTA SLR/SLT model on the dataset and get a baseline SLR word error rate of 34.08% and a baseline SLT BLEU-4 score of 23.58 for benchmarking future research on the dataset.
Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models
Verga, Pat, Hofstatter, Sebastian, Althammer, Sophia, Su, Yixuan, Piktus, Aleksandra, Arkhangorodsky, Arkady, Xu, Minjie, White, Naomi, Lewis, Patrick
As Large Language Models (LLMs) have become more advanced, they have outpaced our abilities to accurately evaluate their quality. Not only is finding data to adequately probe particular model properties difficult, but evaluating the correctness of a model's freeform generation alone is a challenge. To address this, many evaluations now rely on using LLMs themselves as judges to score the quality of outputs from other LLMs. Evaluations most commonly use a single large model like GPT4. While this method has grown in popularity, it is costly, has been shown to introduce intramodel bias, and in this work, we find that very large models are often unnecessary. We propose instead to evaluate models using a Panel of LLm evaluators (PoLL). Across three distinct judge settings and spanning six different datasets, we find that using a PoLL composed of a larger number of smaller models outperforms a single large judge, exhibits less intra-model bias due to its composition of disjoint model families, and does so while being over seven times less expensive.
Eight US newspapers sue OpenAI and Microsoft for copyright infringement
The New York Daily News, Chicago Tribune, Denver Post and other papers filed the lawsuit on Tuesday in a New York federal court. "We've spent billions of dollars gathering information and reporting news at our publications, and we can't allow OpenAI and Microsoft to expand the Big Tech playbook of stealing our work to build their own businesses at our expense," said a written statement from Frank Pine, executive editor for the MediaNews Group and Tribune Publishing. The other newspapers that are part of the lawsuit are MediaNews Group's Mercury News, Denver Post, Orange County Register and St Paul Pioneer-Press, and Tribune Publishing's Orlando Sentinel and South Florida Sun Sentinel. All of the newspapers are owned by Alden Global Capital. Microsoft declined to comment on Tuesday.