Media
HistNERo: Historical Named Entity Recognition for the Romanian Language
Avram, Andrei-Marius, Iuga, Andreea, Manolache, George-Vlad, Matei, Vlad-Cristian, Micliuş, Răzvan-Gabriel, Muntean, Vlad-Andrei, Sorlescu, Manuel-Petru, Şerban, Dragoş-Andrei, Urse, Adrian-Dinu, Păiş, Vasile, Cercel, Dumitru-Clementin
This work introduces HistNERo, the first Romanian corpus for Named Entity Recognition (NER) in historical newspapers. The dataset contains 323k tokens of text, covering more than half of the 19th century (i.e., 1817) until the late part of the 20th century (i.e., 1990). Eight native Romanian speakers annotated the dataset with five named entities. The samples belong to one of the following four historical regions of Romania, namely Bessarabia, Moldavia, Transylvania, and Wallachia. We employed this proposed dataset to perform several experiments for NER using Romanian pre-trained language models. Our results show that the best model achieved a strict F1-score of 55.69%. Also, by reducing the discrepancies between regions through a novel domain adaption technique, we improved the performance on this corpus to a strict F1-score of 66.80%, representing an absolute gain of more than 10%.
Surprising Efficacy of Fine-Tuned Transformers for Fact-Checking over Larger Language Models
In this paper, we explore the challenges associated with establishing an end-to-end fact-checking pipeline in a real-world context, covering over 90 languages. Our real-world experimental benchmarks demonstrate that fine-tuning Transformer models specifically for fact-checking tasks, such as claim detection and veracity prediction, provide superior performance over large language models (LLMs) like GPT-4, GPT-3.5-Turbo, and Mistral-7b. However, we illustrate that LLMs excel in generative tasks such as question decomposition for evidence retrieval. Through extensive evaluation, we show the efficacy of fine-tuned models for fact-checking in a multilingual setting and complex claims that include numerical quantities.
ComposerX: Multi-Agent Symbolic Music Composition with LLMs
Deng, Qixin, Yang, Qikai, Yuan, Ruibin, Huang, Yipeng, Wang, Yi, Liu, Xubo, Tian, Zeyue, Pan, Jiahao, Zhang, Ge, Lin, Hanfeng, Li, Yizhi, Ma, Yinghao, Fu, Jie, Lin, Chenghua, Benetos, Emmanouil, Wang, Wenwu, Xia, Guangyu, Xue, Wei, Guo, Yike
Music composition represents the creative side of humanity, and itself is a complex task that requires abilities to understand and generate information with long dependency and harmony constraints. While demonstrating impressive capabilities in STEM subjects, current LLMs easily fail in this task, generating ill-written music even when equipped with modern techniques like In-Context-Learning and Chain-of-Thoughts. To further explore and enhance LLMs' potential in music composition by leveraging their reasoning ability and the large knowledge base in music history and theory, we propose ComposerX, an agent-based symbolic music generation framework. We find that applying a multi-agent approach significantly improves the music composition quality of GPT-4. The results demonstrate that ComposerX is capable of producing coherent polyphonic music compositions with captivating melodies, while adhering to user instructions.
FactCheck Editor: Multilingual Text Editor with End-to-End fact-checking
We introduce 'FactCheck Editor', an advanced text editor designed to automate fact-checking and correct factual inaccuracies. Given the widespread issue of misinformation, often a result of unintentional mistakes by content creators, our tool aims to address this challenge. It supports over 90 languages and utilizes transformer models to assist humans in the labor-intensive process of fact verification. This demonstration showcases a complete workflow that detects text claims in need of verification, generates relevant search engine queries, and retrieves appropriate documents from the web. It employs Natural Language Inference (NLI) to predict the veracity of claims and uses LLMs to summarize the evidence and suggest textual revisions to correct any errors in the text. Additionally, the effectiveness of models used in claim detection and veracity assessment is evaluated across multiple languages.
Countering Reward Over-optimization in LLM with Demonstration-Guided Reinforcement Learning
Rita, Mathieu, Strub, Florian, Chaabouni, Rahma, Michel, Paul, Dupoux, Emmanuel, Pietquin, Olivier
While Reinforcement Learning (RL) has been proven essential for tuning large language models (LLMs), it can lead to reward over-optimization (ROO). Existing approaches address ROO by adding KL regularization, requiring computationally expensive hyperparameter tuning. Additionally, KL regularization focuses solely on regularizing the language policy, neglecting a potential source of regularization: the reward function itself. Inspired by demonstration-guided RL, we here introduce the Reward Calibration from Demonstration (RCfD), which leverages human demonstrations and a reward model to recalibrate the reward objective. Formally, given a prompt, the RCfD objective minimizes the distance between the demonstrations' and LLM's rewards rather than directly maximizing the reward function. This objective shift avoids incentivizing the LLM to exploit the reward model and promotes more natural and diverse language generation. We show the effectiveness of RCfD on three language tasks, which achieves comparable performance to carefully tuned baselines while mitigating ROO.
OpenAI will train its AI models on the Financial Times' journalism
The Financial Times has become the latest news organization to strike a deal with OpenAI. In a joint announcement on Monday, the Financial Times and OpenAI said that maker of ChatGPT will use the Financial Times' journalism to train its AI models and collaborate on developing new AI products and features for the publication's readers. "It is right, of course, that AI platforms pay publishers for the use of their material," said Financial Times CEO John Ridding in a statement and added that the Times is "committed to human journalism." Neither company disclosed the financial terms of the agreement. Earlier this year, The Information reported that OpenAI offers publishers between 1 million and 5 million a year to license their content to train its AI models.
OpenAI to use FT journalism to train artificial intelligence systems
The Financial Times has struck a deal with ChatGPT developer OpenAI that allows its content to be used in training artificial intelligence systems. The FT will receive an undisclosed payment as part of the deal, which is the latest to be agreed between OpenAI and news publishers. Under the arrangement, ChatGPT users will receive summaries and quotes from FT journalism, as well as links to articles, in responses to prompts, where appropriate. John Ridding, the chief executive of the FT Group, said it was "right" that AI companies paid publishers for their material. The New York Times is suing OpenAI and its largest investor, Microsoft, over use of its content to train large language models, the technology that underpins chatbots like ChatGPT.
The Machine Ethics podcast: Good tech with Eleanor Drage and Kerry McInerney
Hosted by Ben Byford, The Machine Ethics Podcast brings together interviews with academics, authors, business leaders, designers and engineers on the subject of autonomous algorithms, artificial intelligence, machine learning, and technology's impact on society. This episode we're chatting with Eleanor and Kerry on good technology and if it's even possible, that technology is political, watering down regulation, the magic of AI, the value of human creativity, how Feminism, Aboriginal, and mixed race studies can help AI development, the performative nature of tech, and more… Dr Kerry McInerney (née Mackereth) is a Research Fellow at the Leverhulme Centre for the Future of Intelligence at the University of Cambridge, where she co-leads the Global Politics of AI project on how AI is impacting international relations. She is also a Research Fellow at the AI Now Institute (a leading AI policy thinktank in New York), an AHRC/BBC New Generation Thinker (2023), one of the 100 Brilliant Women in AI Ethics (2022), and one of Computing's Rising Stars 30 (2023). Kerry is the co-editor of the collection Feminist AI: Critical Perspectives on Algorithms, Data, and Intelligent Machines (2023, Oxford University Press), the collection The Good Robot: Why Technology Needs Feminism (2024, Bloomsbury Academic), and the co-author of the forthcoming book Reprogram: Why Big Tech is Broken and How Feminism Can Fix It (2026, Princeton University Press). Dr Eleanor Drage is a Senior Research Fellow at the University of Cambridge Centre for the Future of Intelligence, and teaches AI professionals about AI ethics on a Masters course at Cambridge.
Why China Is So Bad at Disinformation
"China will use AI to disrupt elections in the US, South Korea and India, Microsoft warns" one read. "China Is Using AI to Sow Disinformation and Stoke Discord Across Asia and the US," another claimed. The headlines were based on a report published earlier this month by Microsoft's Threat Analysis Center which outlined how a Chinese disinformation campaign was now utilizing artificial technology to inflame divisions and disrupt elections in the US and around the world. The campaign, which has already targeted Taiwan's elections, uses AI-generated audio and memes designed to grab user attention and boost engagement. But what these headlines and Microsoft itself failed to adequately convey is that the Chinese government-linked disinformation campaign, known as Spamouflage Dragon or Dragonbridge, has so far been virtually ineffective.
Point Cloud Models Improve Visual Robustness in Robotic Learners
Peri, Skand, Lee, Iain, Kim, Chanho, Fuxin, Li, Hermans, Tucker, Lee, Stefan
Visual control policies can encounter significant performance degradation when visual conditions like lighting or camera position differ from those seen during training -- often exhibiting sharp declines in capability even for minor differences. In this work, we examine robustness to a suite of these types of visual changes for RGB-D and point cloud based visual control policies. To perform these experiments on both model-free and model-based reinforcement learners, we introduce a novel Point Cloud World Model (PCWM) and point cloud based control policies. Our experiments show that policies that explicitly encode point clouds are significantly more robust than their RGB-D counterparts. Further, we find our proposed PCWM significantly outperforms prior works in terms of sample efficiency during training. Taken together, these results suggest reasoning about the 3D scene through point clouds can improve performance, reduce learning time, and increase robustness for robotic learners. Project Webpage: https://pvskand.github.io/projects/PCWM