Government
Tradeoffs in Streaming Binary Classification under Limited Inspection Resources
Hassanzadeh, Parisa, Dervovic, Danial, Assefa, Samuel, Reddy, Prashant, Veloso, Manuela
Institutions are increasingly relying on machine learning models Given the imbalanced nature of data in this domain, which makes to identify and alert on abnormal events, such as fraud, cyber attacks learning classifiers that efficiently discriminate among the minority and system failures. These alerts often need to be manually and majority class difficult, and the limited resources available investigated by specialists. Given the operational cost of manual inspections, for inspecting time-sensitive risky events, we are interested in understanding the suspicious events are selected by alerting systems with the relationship between the rate of detection from the carefully designed thresholds. In this paper, we consider an imbalanced minority class (i.e., the fraction of samples from the minority class binary classification problem, where events arrive sequentially selected for inspection) and the inspection budget. Specifically, we and only a limited number of suspicious events can be inspected. We focus on applications that involve real-time processing and decisionmaking model the event arrivals as a non-homogeneous Poisson process, and where an abnormal event can only be inspected at the time compare various suspicious event selection methods including those of arrival, and we investigate how different selection policies based based on static and adaptive thresholds. For each method, we analytically on classifier predictions operate in terms of the limited inspection characterize the tradeoff between the minority-class detection budget rather than the decision threshold.
Empowering Local Communities Using Artificial Intelligence
Hsu, Yen-Chia, Huang, Ting-Hao 'Kenneth', Verma, Himanshu, Mauri, Andrea, Nourbakhsh, Illah, Bozzon, Alessandro
Many powerful Artificial Intelligence (AI) techniques have been engineered with the goals of high performance and accuracy. Recently, AI algorithms have been integrated into diverse and real-world applications. It has become an important topic to explore the impact of AI on society from a people-centered perspective. Previous works in citizen science have identified methods of using AI to engage the public in research, such as sustaining participation, verifying data quality, classifying and labeling objects, predicting user interests, and explaining data patterns. These works investigated the challenges regarding how scientists design AI systems for citizens to participate in research projects at a large geographic scale in a generalizable way, such as building applications for citizens globally to participate in completing tasks. In contrast, we are interested in another area that receives significantly less attention: how scientists co-design AI systems "with" local communities to influence a particular geographical region, such as community-based participatory projects. Specifically, this article discusses the challenges of applying AI in Community Citizen Science, a framework to create social impact through community empowerment at an intensely place-based local scale. We provide insights in this under-explored area of focus to connect scientific research closely to social issues and citizen needs.
Data Augmentation Approaches in Natural Language Processing: A Survey
Li, Bohan, Hou, Yutai, Che, Wanxiang
As an effective strategy, data augmentation (DA) alleviates data scarcity scenarios where deep learning techniques may fail. It is widely applied in computer vision then introduced to natural language processing and achieves improvements in many tasks. One of the main focuses of the DA methods is to improve the diversity of training data, thereby helping the model to better generalize to unseen testing data. In this survey, we frame DA methods into three categories based on the diversity of augmented data, including paraphrasing, noising, and sampling. Our paper sets out to analyze DA methods in detail according to the above categories. Further, we also introduce their applications in NLP tasks as well as the challenges.
Formalizing the Generalization-Forgetting Trade-off in Continual Learning
Raghavan, Krishnan, Balaprakash, Prasanna
We formulate the continual learning (CL) problem via dynamic programming and model the trade-off between catastrophic forgetting and generalization as a two-player sequential game. In this approach, player 1 maximizes the cost due to lack of generalization whereas player 2 minimizes the cost due to catastrophic forgetting. We show theoretically that a balance point between the two players exists for each task and that this point is stable (once the balance is achieved, the two players stay at the balance point). Next, we introduce balanced continual learning (BCL), which is designed to attain balance between generalization and forgetting and empirically demonstrate that BCL is comparable to or better than the state of the art.
Unpacking the Black Box: Regulating Algorithmic Decisions
Blattner, Laura, Nelson, Scott, Spiess, Jann
We characterize optimal oversight of algorithms in a world where an agent designs a complex prediction function but a principal is limited in the amount of information she can learn about the prediction function. We show that limiting agents to prediction functions that are simple enough to be fully transparent is inefficient as long as the bias induced by misalignment between principal's and agent's preferences is small relative to the uncertainty about the true state of the world. Algorithmic audits can improve welfare, but the gains depend on the design of the audit tools. Tools that focus on minimizing overall information loss, the focus of many post-hoc explainer tools, will generally be inefficient since they focus on explaining the average behavior of the prediction function rather than sources of mis-prediction, which matter for welfare-relevant outcomes. Targeted tools that focus on the source of incentive misalignment, e.g., excess false positives or racial disparities, can provide first-best solutions. We provide empirical support for our theoretical findings using an application in consumer lending.
Lossy compression of statistical data using quantum annealer
Yoon, Boram, Nguyen, Nga T. T., Chang, Chia Cheng, Rrapaj, Ermal
We present a new lossy compression algorithm for statistical floating-point data through a representation learning with binary variables. The algorithm finds a set of basis vectors and their binary coefficients that precisely reconstruct the original data. The optimization for the basis vectors is performed classically, while binary coefficients are retrieved through both simulated and quantum annealing for comparison. A bias correction procedure is also presented to estimate and eliminate the error and bias introduced from the inexact reconstruction of the lossy compression for statistical data analyses. The compression algorithm is demonstrated on two different datasets of lattice quantum chromodynamics simulations. The results obtained using simulated annealing show 3.5 times better compression performance than the algorithms based on a neural-network autoencoder and principal component analysis. Calculations using quantum annealing also show promising results, but performance is limited by the integrated control error of the quantum processing unit, which yields large uncertainties in the biases and coupling parameters. Hardware comparison is further studied between the previous generation D-Wave 2000Q and the current D-Wave Advantage system. Our study shows that the Advantage system is more likely to obtain low-energy solutions for the problems than the 2000Q.
Breaking down the AI regulations
Companies, governments, and other institutions started embedding artificial intelligence into their products, services, processes, and decision-making to a great extent. This opened great questions on how the data is used by their systems and if any, what are the implications. The answers become even more serious if we take the complex, evolving algorithms that propose health diagnosis, approve a loan, or even autonomously drive a car. Now more than ever, it is essential to develop AI tools that can be trusted and are responsible as AI has and will have wide-ranging economic impacts across manufacturing, transportation, health, education, and many other sectors. This can be done by the development of public sector policies and laws for promoting and regulating AI. It is quite a recent topic among regulators globally as between 2016 and 2020 a wave of AI regulations and guidelines were published in order to maintain social control over the use of algorithms in our everyday lives.
How Are Machine Learning and Artificial Intelligence Used in Cybersecurity?
Artificial intelligence (AI), machine learning (ML), and deep neural networks (DNNs) are the talk of the town these days. However, few people understand the difference between these innovative technologies. Artificial intelligence is an overarching concept that comprises several fields of computer science. It is geared toward solving tasks intrinsic to the human mind, such as speech recognition and object classification. Machine learning is part of the artificial intelligence ecosystem.
AI System Identifies Buildings Damaged by Wildfire
U.S. researchers developed an AI system that helps classifying buildings with wildfire damage by relying solely on post-fire images with 92% accuracy. Wildfires are increasing in frequency and intensity as climate change becomes more pronounced and visible. These are now causing disruptions in urban areas people left homes and their lives behind. Now, they will have to wait anxiously to know the state of their homes and the damage that they will need to fix. Researchers at Stanford University and the California Polytechnic State University (Cal Poly) have developed an Artificial Intelligence (AI) algorithm system called DamageMap that is a damage classifier; it helps identify building damages within minutes of a catastrophe by studying aerial photographs.
The Human Costs of AI
In 2015 a cohort of well-known scientists and entrepreneurs including Stephen Hawking, Elon Musk, and Steve Wozniak issued a public letter urging technologists developing artificial intelligence systems to "research how to reap its benefits while avoiding potential pitfalls." To that end, they wrote, "We recommend expanded research aimed at ensuring that increasingly capable AI systems are robust and beneficial: our AI systems must do what we want them to do." More than eight thousand people have now signed that letter. While most are academics, the signers also include researchers at Palantir, the secretive surveillance firm that helps ICE round up undocumented immigrants; the leaders of Vicarious, an industrial robotics company that boasts reductions for its clients of more than 50 percent in labor hours--which is to say, work performed by humans; and the founders of Sentient Technologies, who had previously developed the language-recognition technology used by Siri, Apple's voice assistant, and whose company has since been folded into Cognizant, a corporation that provided some of the underpaid, overly stressed workforce tasked with "moderating" content on Facebook. Musk, meanwhile, is pursuing more than just AI-equipped self-driving cars. His brain-chip company, Neuralink, aims to merge the brain with artificial intelligence, not only to develop life-changing medical applications for people with spinal cord injuries and neurological disorders, but, eventually, for everyone, to create a kind of hive mind. The goal, according to Musk, is a future "controlled by the combined will of the people of Earth--[since] that's obviously gonna be the future that we want."