Goto

Collaborating Authors

 Government


The Challenges of Machine Learning for Trust and Safety: A Case Study on Misinformation Detection

arXiv.org Artificial Intelligence

We examine the disconnect between scholarship and practice in applying machine learning to trust and safety problems, using misinformation detection as a case study. We systematize literature on automated detection of misinformation across a corpus of 270 well-cited papers in the field. We then examine subsets of papers for data and code availability, design missteps, reproducibility, and generalizability. We find significant shortcomings in the literature that call into question claimed performance and practicality. Detection tasks are often meaningfully distinct from the challenges that online services actually face. Datasets and model evaluation are often non-representative of real-world contexts, and evaluation frequently is not independent of model training. Data and code availability is poor. Models do not generalize well to out-of-domain data. Based on these results, we offer recommendations for evaluating machine learning applications to trust and safety problems. Our aim is for future work to avoid the pitfalls that we identify.


An Open-Source ML-Based Full-Stack Optimization Framework for Machine Learning Accelerators

arXiv.org Artificial Intelligence

Parameterizable machine learning (ML) accelerators are the product of recent breakthroughs in ML. To fully enable their design space exploration (DSE), we propose a physical-design-driven, learning-based prediction framework for hardware-accelerated deep neural network (DNN) and non-DNN ML algorithms. It adopts a unified approach that combines backend power, performance, and area (PPA) analysis with frontend performance simulation, thereby achieving a realistic estimation of both backend PPA and system metrics such as runtime and energy. In addition, our framework includes a fully automated DSE technique, which optimizes backend and system metrics through an automated search of architectural and backend parameters. Experimental studies show that our approach consistently predicts backend PPA and system metrics with an average 7% or less prediction error for the ASIC implementation of two deep learning accelerator platforms, VTA and VeriGOOD-ML, in both a commercial 12 nm process and a research-oriented 45 nm process.


Layer-wise Feedback Propagation

arXiv.org Artificial Intelligence

In this paper, we present Layer-wise Feedback Propagation (LFP), a novel training approach for neural-network-like predictors that utilizes explainability, specifically Layer-wise Relevance Propagation(LRP), to assign rewards to individual connections based on their respective contributions to solving a given task. This differs from traditional gradient descent, which updates parameters towards anestimated loss minimum. LFP distributes a reward signal throughout the model without the need for gradient computations. It then strengthens structures that receive positive feedback while reducingthe influence of structures that receive negative feedback. We establish the convergence of LFP theoretically and empirically, and demonstrate its effectiveness in achieving comparable performance to gradient descent on various models and datasets. Notably, LFP overcomes certain limitations associated with gradient-based methods, such as reliance on meaningful derivatives. We further investigate how the different LRP-rules can be extended to LFP, what their effects are on training, as well as potential applications, such as training models with no meaningful derivatives, e.g., step-function activated Spiking Neural Networks (SNNs), or for transfer learning, to efficiently utilize existing knowledge.


A Massive Scale Semantic Similarity Dataset of Historical English

arXiv.org Artificial Intelligence

A diversity of tasks use language models trained on semantic similarity data. While there are a variety of datasets that capture semantic similarity, they are either constructed from modern web data or are relatively small datasets created in the past decade by human annotators. This study utilizes a novel source, newly digitized articles from off-copyright, local U.S. newspapers, to assemble a massive-scale semantic similarity dataset spanning 70 years from 1920 to 1989 and containing nearly 400M positive semantic similarity pairs. Historically, around half of articles in U.S. local newspapers came from newswires like the Associated Press. While local papers reproduced articles from the newswire, they wrote their own headlines, which form abstractive summaries of the associated articles. We associate articles and their headlines by exploiting document layouts and language understanding. We then use deep neural methods to detect which articles are from the same underlying source, in the presence of substantial noise and abridgement. The headlines of reproduced articles form positive semantic similarity pairs. The resulting publicly available HEADLINES dataset is significantly larger than most existing semantic similarity datasets and covers a much longer span of time. It will facilitate the application of contrastively trained semantic similarity models to a variety of tasks, including the study of semantic change across space and time.


The (Computational) Social Choice Take on Indivisible Participatory Budgeting

arXiv.org Artificial Intelligence

In this survey, we review the literature investigating participatory budgeting as a social choice problem. Participatory Budgeting (PB) is a democratic process in which citizens are asked to vote on how to allocate a given amount of public money to a set of projects. From a social choice perspective, it corresponds then to the problem of aggregating opinions about which projects should be funded, into a budget allocation satisfying a budget constraint. This problem has received substantial attention in recent years and the literature is growing at a fast pace. In this survey, we present the most important research directions from the literature, each time presenting a large set of representative results. We only focus on the indivisible case, that is, PB problems in which projects can either be fully funded or not at all. The aim of the survey is to present a comprehensive overview of the state of the research on PB. We aim at providing both a general overview of the main research questions that are being investigated, and formal and unified definitions of the most important technical concepts from the literature.


A Rigorous Uncertainty-Aware Quantification Framework Is Essential for Reproducible and Replicable Machine Learning Workflows

arXiv.org Artificial Intelligence

The ability to replicate predictions by machine learning (ML) or artificial intelligence (AI) models and results in scientific workflows that incorporate such ML/AI predictions is driven by numerous factors. An uncertainty-aware metric that can quantitatively assess the reproducibility of quantities of interest (QoI) would contribute to the trustworthiness of results obtained from scientific workflows involving ML/AI models. In this article, we discuss how uncertainty quantification (UQ) in a Bayesian paradigm can provide a general and rigorous framework for quantifying reproducibility for complex scientific workflows. Such as framework has the potential to fill a critical gap that currently exists in ML/AI for scientific workflows, as it will enable researchers to determine the impact of ML/AI model prediction variability on the predictive outcomes of ML/AI-powered workflows. We expect that the envisioned framework will contribute to the design of more reproducible and trustworthy workflows for diverse scientific applications, and ultimately, accelerate scientific discoveries.


LAPD to use AI to analyze body cam videos for officers' language use

Los Angeles Times

Researchers will use artificial intelligence to analyze the tone and word choice that LAPD officers use during traffic stops, the department announced Tuesday, part of a broader study of whether police language sometimes unnecessarily escalates public encounters. Findings from the study, conducted by researchers from USC and elsewhere, will be used to help train officers on how best to navigate encounters with the public and to "promote accountability," said Cmdr. Machine learning, she said at a meeting of the Board of Police Commissioners, "is in its infancy, but will undoubtedly become a profound element in officer training in the future." Over three years, researchers will review body camera footage from roughly 1,000 traffic stops, then develop criteria on what constitutes an appropriate interaction based on public and office feedback and a review of the department's policies, according to Benjamin A.T. Graham, an associate professor of international relations at USC and one of the study's authors. These criteria will then be fed into a machine learning program, which will "learn" how to review videos on its own and flag instances where officers cross the line, Graham said.


Iran unveils armed drone resembling America's MQ-9 Reaper, claims it could reach Israel

FOX News

Fox News Flash top headlines are here. Check out what's clicking on Foxnews.com. Iran on Tuesday unveiled a new armed drone bearing resemblance to America's MQ-9 Reaper, with state media claiming it has the operational range to reach Israel. The Mohajer-10 was showcased during a ceremony celebrating the Islamic Republic's Defense Industry Day. The drone can fly non-stop for 24 hours with an operational range of around 1,200 miles and is capable of carrying a bomb payload of up to 660 pounds, according to the state-run IRNA News Agency.


Iran unveils attack drone capable of striking Israel

Al Jazeera

Tehran, Iran – Iran has unveiled a new drone that it says is capable of striking targets in Israel. The Iranian Ministry of Defence and Armed Forces Logistics unveiled the Mohajer-10 on Tuesday as part of an exhibition and ceremonies marking Defence Industry Day. President Ebrahim Raisi and senior commanders in the army and the Islamic Revolutionary Guard Corps (IRGC) attended the event. The unmanned attack aircraft, which resembles the MQ-9 Reaper manufactured by the United States, was also shown in videos taking off from an unidentified airstrip and flying. It is said to be capable of carrying a variety of bombs and anti-radar equipment and of conducting surveillance.


Disasters like Hilary are a magnet for fraudsters. Look out for these scams

Los Angeles Times

If you live in the path of Tropical Storm Hilary, you may soon hear from people offering you the help you desperately need. And the rest of us are probably already being contacted by organizations raising money for disaster victims here, in Maui and in other wildfire-wracked communities. Unfortunately, some of the outreach will come from people who want to harm, not help. Online scammers are eager to take advantage of people's need for aid, as well as their desire to help. Using the internet's ability to conceal their true identity, they pose as government officials, charities and community groups as they try to collect money or sensitive personal information they can sell online.