Goto

Collaborating Authors

 Government


Achieving Equalized Odds by Resampling Sensitive Attributes

arXiv.org Machine Learning

We present a flexible framework for learning predictive models that approximately satisfy the equalized odds notion of fairness. This is achieved by introducing a general discrepancy functional that rigorously quantifies violations of this criterion. This differentiable functional is used as a penalty driving the model parameters towards equalized odds. To rigorously evaluate fitted models, we develop a formal hypothesis test to detect whether a prediction rule violates this property, the first such test in the literature. Both the model fitting and hypothesis testing leverage a resampled version of the sensitive attribute obeying equalized odds, by construction. We demonstrate the applicability and validity of the proposed framework both in regression and multi-class classification problems, reporting improved performance over state-of-the-art methods. Lastly, we show how to incorporate techniques for equitable uncertainty quantification---unbiased for each group under study---to communicate the results of the data analysis in exact terms.


Uncertainty-Aware Deep Classifiers using Generative Models

arXiv.org Machine Learning

Deep neural networks are often ignorant about what they do not know and overconfident when they make uninformed predictions. Some recent approaches quantify classification uncertainty directly by training the model to output high uncertainty for the data samples close to class boundaries or from the outside of the training distribution. These approaches use an auxiliary data set during training to represent out-of-distribution samples. However, selection or creation of such an auxiliary data set is non-trivial, especially for high dimensional data such as images. In this work we develop a novel neural network model that is able to express both aleatoric and epistemic uncertainty to distinguish decision boundary and out-of-distribution regions of the feature space. To this end, variational autoencoders and generative adversarial networks are incorporated to automatically generate out-of-distribution exemplars for training. Through extensive analysis, we demonstrate that the proposed approach provides better estimates of uncertainty for in- and out-of-distribution samples, and adversarial examples on well-known data sets against state-of-the-art approaches including recent Bayesian approaches for neural networks and anomaly detection methods.


Cybersecurity with Data Science: Implementation

#artificialintelligence

Hands-on of Machine Learning in Cybersecurity Supervised and unsupervised machine learning models for cybersecurity Description Machine learning is disrupting cybersecurity to a greater extent than almost any other industry. Many problems in cyber security are well suited to the application of machine learning as they often involve some form of anomaly detection on very large volumes of data. This course deals the most found issues in cybersecurity such as malware, anomalies detection, SQL injection, credit card fraud, bots, spams and phishing. All these problems are covered in case studies.


Elon Musk: Everyone, "including Tesla," needs AI regulation

#artificialintelligence

Musk was responding to a massive feature story published in the MIT Technology Review about OpenAI, the AI research lab founded in part by Elon Musk, alongside others. The lab operates with the mission of developing safe and ethical AI that'll be good for the world. But MIT Tech's reporting tells of how Open AI went from being a transparent organization to a relatively opaque one (hence Musk's preceding Tweet about OpenAI needing to "be more open"). Musk's ability to self-aggrandize or self-flagellate is usually surprising in equal measure, but never shocking: Industries often argue for their own regulation as a way to keep government regulators off their backs. Though credit where it's due: Musk has been, as in the case of when he argued in favor of regulating autonomous weapons, more substantially -- and more effectively -- vocal than most when it comes to regulating AI. Whether or not this will have any substantial effects on other companies (statements from CEOs, regulatory commission efforts, etc) let alone Tesla or OpenAI will be nothing if not a compelling plot to watch.


Lessons in Rapid Innovation From the COVID-19 Pandemic

#artificialintelligence

The coronavirus pandemic is one of the most difficult collective challenges facing humanity since the last world war. In the midst of the turmoil, national health authorities, pharmaceutical companies, universities, and research institutes are racing to find therapies to save lives and contain the grave social and economic consequences of the pandemic. As organizations and experts scramble to innovate therapies, they are also redefining innovation. The conventional approach to innovation in the pharmaceutical industry is to conduct a lengthy process that starts with the discovery and generation of potential drug compounds and moves through a meticulous refinement and selection phase toward gradual development, clinical testing, and market approval. Although this model will continue to be the most effective in future drug development, it is now being complemented with an ultrafast approach to innovation centered on the repurposing of readily available ideas, knowledge, and technologies.


How to protect your identity while protesting police brutality

Engadget

The response to protests against police brutality, ignited by the murder of Geoge Floyd, have been nothing short of draconian. While government forces on the ground gleefully beat protesters and passersby with batons and doused them with tear gas, the US Border Patrol has deployed Reaper drones to surveil citizens from the skies and the DEA has been tasked with tracking protesters. The outsized surveillance response displayed so far by the Feds has driven concerns from privacy advocates over the potential use of more insidious forms of snooping, from facial recognition algorithms to cell-site simulation (aka the Stingray and Crossbow systems.) People stuck in traffic are witnessing NYPD beat up folks on their way home. "All the technology we have been warning about for a while are starting to come to fruition in these protests," Dave Maass, a senior investigative researcher at digital rights group the Electronic Frontier Foundation, told Reuters on Monday.


An Efficient $k$-modes Algorithm for Clustering Categorical Datasets

arXiv.org Machine Learning

Mining clusters from datasets is an important endeavor in many applications. The k-means algorithm is a popular and efficient, distribution-free approach for clustering numerical-valued data, but does not apply for categorical-valued observations. We provide a novel, computationally efficient implementation of k-modes, called OTQT. We prove that OTQT finds updates, undetectable to existing k-modes algorithms, that improve the objective function. Thus, although slightly slower per iteration owing to its algorithmic complexity, OTQT is always more accurate per iteration and almost always faster (and only barely slower on some datasets) to the final optimum. As a result, we recommend OTQT as the preferred, default algorithm for all k-modes implementations. We also examine five initialization methods and three types of K-selection methods, many of them novel or novel applications to k-modes. By examining performance on real and simulated datasets, we show that simple random initialization is the best initializer and that a novel K-selection method is more accurate than methods adapted from k-means. Identifying groups of similar observations in datasets is common in a wide array of applications, with many clustering methods developed in statistics, machine learning and the applied sciences [1]-[7]. The k-means algorithm [8]-[11] is arguably the most popular method for clustering numerical-valued observations. It scales to large datasets because it does not require calculation of all pairwise distances, and it is distribution-free. While distribution-free does not imply it is assumption-free [12], [13], it is a starting place for users wary of making assumptions about their data. Unfortunately, k-means does not provide an appropriate objective to minimize for datasets with categorical attributes.


Deepfakes Are Going To Wreak Havoc On Society. We Are Not Prepared.

#artificialintelligence

None of these people exist. These images were generated using deepfake technology. Last month during ESPN's hit documentary series The Last Dance, State Farm debuted a TV commercial that has become one of the most widely discussed ads in recent memory. It appeared to show footage from 1998 of an ESPN analyst making shockingly accurate predictions about the year 2020. As it turned out, the clip was not genuine: it was generated using cutting-edge AI.


Nuclear Fusion and Artificial Intelligence: the Dream of Limitless Energy

#artificialintelligence

Ever since the 1930s when scientists, namely Hans Bethe, discovered that nuclear fusion was possible, researchers strived to initiate and control fusion reactions to produce useful energy on Earth. The best example of a fusion reaction is in the middle of stars like the Sun where hydrogen atoms are fused together to make helium releasing a lot of energy that powers the heat and light of the star. On Earth, scientists need to heat and control plasma, an ionised state of matter similar to gas, to cause particles to fuse and release their energy. Unfortunately, it is very difficult to start fusion reactions on Earth, as they require conditions similar to the Sun, very high temperature and pressure, and scientists have been trying to find a solution for decades. In May 2019, a workshop detailing how fusion could be advanced using machine learning was held that was jointly supported by the Department of Energy Offices of Fusion Energy Science (FES) and Advanced Scientific Computing Research (ASCR).


The ultimate root of AI stupidity

#artificialintelligence

This is the second installment in a series. The ultimate root of the stupidity of AI systems, I argue, lies in their strictly algorithmic character. AI as presently understood is based on digital processing systems that carry out binary-numerical operations in a step-by-step fashion according to fixed sets of algorithms, starting from an array of numerical inputs. Some may object to this characterization, pointing out that AI systems can constantly change their own "rules" – reprogramming themselves, so to speak. That is true; but the self-reprogramming must follow some algorithm.