Government
IAI Group at CheckThat! 2024: Transformer Models and Data Augmentation for Checkworthy Claim Detection
Aarnes, Peter Røysland, Setty, Vinay, Galuščáková, Petra
This paper describes IAI group's participation for automated check-worthiness estimation for claims, within the framework of the 2024 CheckThat! Lab "Task 1: Check-Worthiness Estimation". The task involves the automated detection of check-worthy claims in English, Dutch, and Arabic political debates and Twitter data. We utilized various pre-trained generative decoder and encoder transformer models, employing methods such as few-shot chain-of-thought reasoning, fine-tuning, data augmentation, and transfer learning from one language to another. Despite variable success in terms of performance, our models achieved notable placements on the organizer's leaderboard: ninth-best in English, third-best in Dutch, and the top placement in Arabic, utilizing multilingual datasets for enhancing the generalizability of check-worthiness detection. Despite a significant drop in performance on the unlabeled test dataset compared to the development test dataset, our findings contribute to the ongoing efforts in claim detection research, highlighting the challenges and potential of language-specific adaptations in claim verification systems.
Detection and Characterization of Coordinated Online Behavior: A Survey
Mannocci, Lorenzo, Mazza, Michele, Monreale, Anna, Tesconi, Maurizio, Cresci, Stefano
Coordination is a fundamental aspect of life. The advent of social media has made it integral also to online human interactions, such as those that characterize thriving online communities and social movements. At the same time, coordination is also core to effective disinformation, manipulation, and hate campaigns. This survey collects, categorizes, and critically discusses the body of work produced as a result of the growing interest on coordinated online behavior. We reconcile industry and academic definitions, propose a comprehensive framework to study coordinated online behavior, and review and critically discuss the existing detection and characterization methods. Our analysis identifies open challenges and promising directions of research, serving as a guide for scholars, practitioners, and policymakers in understanding and addressing the complexities inherent to online coordination.
DebateQA: Evaluating Question Answering on Debatable Knowledge
Xu, Rongwu, Qi, Xuan, Qi, Zehan, Xu, Wei, Guo, Zhijiang
The rise of large language models (LLMs) has enabled us to seek answers to inherently debatable questions on LLM chatbots, necessitating a reliable way to evaluate their ability. However, traditional QA benchmarks assume fixed answers are inadequate for this purpose. To address this, we introduce DebateQA, a dataset of 2,941 debatable questions, each accompanied by multiple human-annotated partial answers that capture a variety of perspectives. We develop two metrics: Perspective Diversity, which evaluates the comprehensiveness of perspectives, and Dispute Awareness, which assesses if the LLM acknowledges the question's debatable nature. Experiments demonstrate that both metrics align with human preferences and are stable across different underlying models. Using DebateQA with two metrics, we assess 12 popular LLMs and retrieval-augmented generation methods. Our findings reveal that while LLMs generally excel at recognizing debatable issues, their ability to provide comprehensive answers encompassing diverse perspectives varies considerably.
Reconsidering Token Embeddings with the Definitions for Pre-trained Language Models
Zhang, Ying, Li, Dongyuan, Okumura, Manabu
Learning token embeddings based on token co-occurrence statistics has proven effective for both pre-training and fine-tuning in natural language processing. However, recent studies have pointed out the distribution of learned embeddings degenerates into anisotropy, and even pre-trained language models (PLMs) suffer from a loss of semantics-related information in embeddings for low-frequency tokens. This study first analyzes fine-tuning dynamics of a PLM, BART-large, and demonstrates its robustness against degeneration. On the basis of this finding, we propose DefinitionEMB, a method that utilizes definitions to construct isotropically distributed and semantics-related token embeddings for PLMs while maintaining original robustness during fine-tuning. Our experiments demonstrate the effectiveness of leveraging definitions from Wiktionary to construct such embeddings for RoBERTa-base and BART-large. Furthermore, the constructed embeddings for low-frequency tokens improve the performance of these models across various GLUE and four text summarization datasets.
"A Good Bot Always Knows Its Limitations": Assessing Autonomous System Decision-making Competencies through Factorized Machine Self-confidence
Israelsen, Brett, Ahmed, Nisar R., Aitken, Matthew, Frew, Eric W., Lawrence, Dale A., Argrow, Brian M.
How can intelligent machines assess their competencies in completing tasks? This question has come into focus for autonomous systems that algorithmically reason and make decisions under uncertainty. It is argued here that machine self-confidence - a form of meta-reasoning based on self-assessments of an agent's knowledge about the state of the world and itself, as well as its ability to reason about and execute tasks - leads to many eminently computable and useful competency indicators for such agents. This paper presents a culmination of work on this concept in the form of a computational framework called Factorized Machine Self-confidence (FaMSeC), which provides a holistic engineering-focused description of factors driving an algorithmic decision-making process, including: outcome assessment, solver quality, model quality, alignment quality, and past experience. In FaMSeC, self confidence indicators are derived from hierarchical `problem-solving statistics' embedded within broad classes of probabilistic decision-making algorithms such as Markov decision processes. The problem-solving statistics are obtained by evaluating and grading probabilistic exceedance margins with respect to given competency standards, which are specified for each of the various decision-making competency factors by the informee (e.g. a non-expert user or an expert system designer). This approach allows `algorithmic goodness of fit' evaluations to be easily incorporated into the design of many kinds of autonomous agents in the form of human-interpretable competency self-assessment reports. Detailed descriptions and application examples for a Markov decision process agent show how two of the FaMSeC factors (outcome assessment and solver quality) can be computed and reported for a range of possible tasking contexts through novel use of meta-utility functions, behavior simulations, and surrogate prediction models.
Government shelves 1.3bn UK tech and AI plans
The Conservatives said that under its leadership, the department had underspent. Those affected have been notified by Secretary of State Peter Kyle. "The government is taking difficult and necessary spending decisions across all departments in the face of billions of pounds of unfunded commitments," said DSIT in a statement. "This is essential to restore economic stability and deliver our national mission for growth." It added that it remained "absolutely committed" to building technology infrastructure in the UK.
Real-life Minority Report: Argentina will use AI to 'predict future crimes'
Argentinian security forces have announced plans to use artificial intelligence to'predict future crimes' but experts warn the move could threaten citizens' rights. Far-right president Javier Milei has created the Artificial Intelligence Applied to Security Unit which will use algorithms to analyse historical crime data. The data produced will then be used to predict future crimes, The Guardian has reported. The security unit is also expected to be able to use facial recognition software to track down wanted persons and detect suspicious activity. However, the Minority Report-esque resolution has concerned human rights campaigners who fear certain groups in society may be over-scrutinised by the AI technology.
The Morning After: Squid Game returns on December 26
After the live experiences, TV shows based on TV shows and a boom in childhood South Korean games and hobbies, Squid Game returns for season two. Almost three years after the bleak, lightly anti-capitalism drama became a massive hit in the US. Season two will hit Netflix December 26, with a final third season coming sometime in 2025. In a letter, series director and writer, Hwang Dong-hyuk, teased the continuation of Seong Gi-hun's revenge, facing off against Front Man. We're expecting more death, betrayal and enough delicious Korean food to make me want to take a trip to Seoul.
Russian advances in Donetsk threaten Ukrainian lines of supply
During the last week of July, Russia mounted its largest assaults in eight months in eastern Ukraine's Donetsk region, seizing a string of settlements in an apparent bid to cut off key supply routes and force a mass Ukrainian retreat. At the same time, Ukraine scored a high number of hits on Russian energy infrastructure and occupied Crimea, suggesting that its strategy of degrading Russian air defences is working. Russian assaults focused on central and southern Donetsk – from areas west of Bakhmut, which fell in May last year, to areas west of Avdiivka, which was lost in February, down to areas west of the city of Donetsk, which pro-Moscow separatists have controlled since 2014 – a line about 130km (80 miles) long. Russian forces have pressed their advantage in these areas to prevent Ukraine from digging entrenched defences, and they have inched forward for months, swallowing settlements at a staggering cost to their own troops. British military intelligence estimated that Russian casualties in May and June reached record daily highs of about 1,200 – about 70,000 soldiers for just those two months.
OpenAI vows to provide the US government early access to its next AI model
OpenAI will give the US AI Safety Institute early access to its next model as part of its safety efforts, Sam Altman has revealed in a tweet. Apparently, the company has been working with the consortium "to push forward the science of AI evaluations." The National Institute of Standards and Technology (NIST) has formally established the Artificial Intelligence Safety Institute earlier this year, though Vice President Kamala Harris announced it back in 2023 at the UK AI Safety Summit. Based on the NIST's description of the consortium, it's meant "to develop science-based and empirically backed guidelines and standards for AI measurement and policy, laying the foundation for AI safety across the world." The company, along with DeepMind, similarly pledged to share AI models with the UK government last year.