Government
Implications of Distance over Redistricting Maps: Central and Outlier Maps
Esmaeili, Seyed A., Chakrabarti, Darshan, Grape, Hayley, Brubach, Brian
In representative democracy, a redistricting map is chosen to partition an electorate into a collection of districts each of which elects a representative. A valid redistricting map must satisfy a collection of constraints such as being compact, contiguous, and of almost equal population. However, these imposed constraints are still loose enough to enable an enormous ensemble of valid redistricting maps. This fact introduces a difficulty in drawing redistricting maps and it also enables a partisan legislature to possibly gerrymander by choosing a map which unfairly favors it. In this paper, we introduce an interpretable and tractable distance measure over redistricting maps which does not use election results and study its implications over the ensemble of redistricting maps. Specifically, we define a central map which may be considered as being "most typical" and give a rigorous justification for it by showing that it mirrors the Kemeny ranking in a scenario where we have a committee voting over a collection of redistricting maps to be drawn. We include run-time and sample complexity analysis for our algorithms, including some negative results which hold using any algorithm. We further study outlier detection based on this distance measure. More precisely, we show gerrymandered maps that lie very far away from our central maps in comparison to a large ensemble of valid redistricting maps. Since our distance measure does not rely on election results, this gives a significant advantage in gerrymandering detection which is lacking in all previous methods.
Adversarial Attacks on Online Learning to Rank with Stochastic Click Models
Wang, Zichen, Balasubramanian, Rishab, Yuan, Hui, Song, Chenyu, Wang, Mengdi, Wang, Huazheng
Online learning to rank (OLTR) (Grotov and de Rijke, 2016) formulates learning to rank (Liu et al., 2009), the core problem in information retrieval, as a sequential decision-making problem. OLTR is a family of online learning solutions that exploit implicit feedback from users (e.g., clicks) to directly optimize parameterized rankers on the fly. It has drawn increasing attention in recent years (Kveton et al., 2015a; Zoghi et al., 2017; Lattimore et al., 2018; Oosterhuis and de Rijke, 2018; Wang et al., 2019; Jia et al., 2021) due to its advantages over traditional offline learning-based solutions and numerous applications in web search and recommender systems (Liu et al., 2009). To effectively utilize users' click feedback to improve the quality of ranked lists, one line of OLTR studied bandit-based algorithms under different click models. In each iteration, the algorithm presents a ranked list of K items selected from L candidates based on its estimation of the user's interests. The ranker observes the user's click feedback and updates these estimates accordingly. Different users may examine and click on the ranking list differently, and how the user interacts with the item list is called the click model. Many works have been dedicated to establishing OLTR algorithms in the cascade model (Kveton et al., 2015a,b; Zong et al., 2016; Li et al., 2016; Vial et al.,
Design and implementation of intelligent packet filtering in IoT microcontroller-based devices
Bertoli, Gustavo de Carvalho, Fernandes, Gabriel Victor C., Monici, Pedro H. Borges, Guibo, Cรฉsar H. de Araujo, Pereira, Lourenรงo Alves Jr., Santos, Aldri
Internet of Things (IoT) devices are increasingly pervasive and essential components in enabling new applications and services. However, their widespread use also exposes them to exploitable vulnerabilities and flaws that can lead to significant losses. In this context, ensuring robust cybersecurity measures is essential to protect IoT devices from malicious attacks. However, the current solutions that provide flexible policy specifications and higher security levels for IoT devices are scarce. To address this gap, we introduce T800, a low-resource packet filter that utilizes machine learning (ML) algorithms to classify packets in IoT devices. We present a detailed performance benchmarking framework and demonstrate T800's effectiveness on the ESP32 system-on-chip microcontroller and ESP-IDF framework. Our evaluation shows that T800 is an efficient solution that increases device computational capacity by excluding unsolicited malicious traffic from the processing pipeline. Additionally, T800 is adaptable to different systems and provides a well-documented performance evaluation strategy for security ML-based mechanisms on ESP32-based IoT systems. Our research contributes to improving the cybersecurity of resource-constrained IoT devices and provides a scalable, efficient solution that can be used to enhance the security of IoT systems.
Shapley Based Residual Decomposition for Instance Analysis
In this paper, we introduce the idea of decomposing the residuals of regression with respect to the data instances instead of features. This allows us to determine the effects of each individual instance on the model and each other, and in doing so makes for a model-agnostic method of identifying instances of interest. In doing so, we can also determine the appropriateness of the model and data in the wider context of a given study. The paper focuses on the possible applications that such a framework brings to the relatively unexplored field of instance analysis in the context of Explainable AI tasks.
Leveraging Domain Knowledge for Inclusive and Bias-aware Humanitarian Response Entry Classification
Tamagnone, Nicolรฒ, Fekih, Selim, Contla, Ximena, Orozco, Nayid, Rekabsaz, Navid
Accurate and rapid situation analysis during humanitarian crises is critical to delivering humanitarian aid efficiently and is fundamental to humanitarian imperatives and the Leave No One Behind (LNOB) principle. This data analysis can highly benefit from language processing systems, e.g., by classifying the text data according to a humanitarian ontology. However, approaching this by simply fine-tuning a generic large language model (LLM) involves considerable practical and ethical issues, particularly the lack of effectiveness on data-sparse and complex subdomains, and the encoding of societal biases and unwanted associations. In this work, we aim to provide an effective and ethically-aware system for humanitarian data analysis. We approach this by (1) introducing a novel architecture adjusted to the humanitarian analysis framework, (2) creating and releasing a novel humanitarian-specific LLM called HumBert, and (3) proposing a systematic way to measure and mitigate biases. Our experiments' results show the better performance of our approach on zero-shot and full-training settings in comparison with strong baseline models, while also revealing the existence of biases in the resulting LLMs. Utilizing a targeted counterfactual data augmentation approach, we significantly reduce these biases without compromising performance.
Semantically-informed Hierarchical Event Modeling
Dipta, Shubhashis Roy, Rezaee, Mehdi, Ferraro, Francis
Prior work has shown that coupling sequential latent variable models with semantic ontological knowledge can improve the representational capabilities of event modeling approaches. In this work, we present a novel, doubly hierarchical, semi-supervised event modeling framework that provides structural hierarchy while also accounting for ontological hierarchy. Our approach consists of multiple layers of structured latent variables, where each successive layer compresses and abstracts the previous layers. We guide this compression through the injection of structured ontological knowledge that is defined at the type level of events: importantly, our model allows for partial injection of semantic knowledge and it does not depend on observing instances at any particular level of the semantic ontology. Across two different datasets and four different evaluation metrics, we demonstrate that our approach is able to out-perform the previous state-of-the-art approaches by up to 8.5%, demonstrating the benefits of structured and semantic hierarchical knowledge for event modeling.
China to land astronauts on moon before 2030, officials say
Former NASA astronaut Tom Jones speaks on what the launch of Artemis I could mean for the future of space exploration on'Your World.' China space officials said Monday that the program plans to place astronauts on the moon before 2030, as well as expand its space station. The deputy director of the Chinese Manned Space Agency confirmed that objectives at a press conference at the Jiuquan Satellite Launch Center, but did not provide a timeline. Deputy Director Lin Xiqiang told reporters that the country is first preparing for a "short stay on the lunar surface and human-robotic joint exploration." "We have a complete near-Earth human space station and human round-trip transportation system," he said.
Russia issues Lindsey Graham arrest warrant after Ukraine comments
President of Ukraine Volodymyr Zelenskyy held a meeting with U.S. Sen. Lindsey Graham on May 26, 2023, during the senator's third visit to Ukraine since Russia invaded the country. Sen. Lindsey Graham, R-S.C., is a wanted man in Russia for comments he made while visiting Ukrainian President Volodymyr Zelenskyy on Friday. Russia's Interior Ministry put out a warrant for Graham's arrest on Monday in response to an edited video released by Zelenskyy's office in which Graham praised U.S. support for Ukraine's defense and noted that Russians are dying as Ukraine fights for its freedom. In the video, Graham noted that "the Russians are dying" and described the U.S. military assistance to the country as "the best money we've ever spent." While Graham appeared to have made the remarks in different parts of the conversation, the short video by Ukraine's presidential office put them next to each other, causing outrage in Russia.
AI-generated misinformation likely to pose hazard in U.S. election campaigns
Fast-evolving AI technology could turbocharge misinformation in U.S. political campaigns, observers say. The 2024 presidential race is expected to be the first American election that will see the widespread use of advanced tools powered by artificial intelligence that have increasingly blurred the boundaries between fact and fiction. Campaigns on both sides of the political divide are likely to harness this technology -- which is cheap, easily accessible and whose advances have vastly outpaced regulatory responses -- for voter outreach and to churn out fundraising newsletters within seconds. This could be due to a conflict with your ad-blocking or security software. Please add japantimes.co.jp and piano.io to your list of allowed sites.
Extractive is not Faithful: An Investigation of Broad Unfaithfulness Problems in Extractive Summarization
Zhang, Shiyue, Wan, David, Bansal, Mohit
The problems of unfaithful summaries have been widely discussed under the context of abstractive summarization. Though extractive summarization is less prone to the common unfaithfulness issues of abstractive summaries, does that mean extractive is equal to faithful? Turns out that the answer is no. In this work, we define a typology with five types of broad unfaithfulness problems (including and beyond not-entailment) that can appear in extractive summaries, including incorrect coreference, incomplete coreference, incorrect discourse, incomplete discourse, as well as other misleading information. We ask humans to label these problems out of 1600 English summaries produced by 16 diverse extractive systems. We find that 30% of the summaries have at least one of the five issues. To automatically detect these problems, we find that 5 existing faithfulness evaluation metrics for summarization have poor correlations with human judgment. To remedy this, we propose a new metric, ExtEval, that is designed for detecting unfaithful extractive summaries and is shown to have the best performance. We hope our work can increase the awareness of unfaithfulness problems in extractive summarization and help future work to evaluate and resolve these issues. Our data and code are publicly available at https://github.com/ZhangShiyue/extractive_is_not_faithful