Government
Conquering AI risks
The age of pervasive AI is here.1 Since 2017, Deloitte's annual State of AI in the Enterprise report has measured the rapid advancement of AI technology globally and across industries. In the most recent edition, published in July 2020, a majority of those surveyed reported significant increases in AI investments, with more than three-quarters believing that AI will substantially transform their organization in the next three years. In addition, AI investments are increasingly leading to measurable organizational benefits: improved process efficiency, better decision-making, increased worker productivity, and enhanced products and services.2 These possible benefits have likely driven the growth in AI's perceived value to organizations--nearly three-quarters of respondents report that AI is strategically important, an increase of 10 percentage points from the previous survey.
Stealth (film) - Wikipedia
Stealth is a 2005 American military science fiction action film directed by Rob Cohen and written by W. D. Richter, and starring Josh Lucas, Jessica Biel, Jamie Foxx, Sam Shepard, Joe Morton and Richard Roxburgh. The film follows three top fighter pilots as they join a project to develop an automated robotic stealth aircraft. Released on July 29, 2005 by Columbia Pictures, the film was a box office bomb, grossing $79 million worldwide against a budget of $135 million. It was one of the worst losses in cinematic history.[2][3] In the near future, the U.S. Navy develops the F/A-37 Talon, a single-seat fighter-bomber with advanced payload, range, speed, and stealth capabilities.
Dual regularized Laplacian spectral clustering methods on community detection
Spectral clustering methods are widely used for detecting clusters in networks for community detection, while a small change on the graph Laplacian matrix could bring a dramatic improvement. In this paper, we propose a dual regularized graph Laplacian matrix and then employ it to three classical spectral clustering approaches under the degree-corrected stochastic block model. If the number of communities is known as $K$, we consider more than $K$ leading eigenvectors and weight them by their corresponding eigenvalues in the spectral clustering procedure to improve the performance. Three improved spectral clustering methods are dual regularized spectral clustering (DRSC) method, dual regularized spectral clustering on Ratios-of-eigenvectors (DRSCORE) method, and dual regularized symmetrized Laplacian inverse matrix (DRSLIM) method. Theoretical analysis of DRSC and DRSLIM show that under mild conditions DRSC and DRSLIM yield stable consistent community detection, moreover, DRSCORE returns perfect clustering under the ideal case. We compare the performances of DRSC, DRSCORE and DRSLIM with several spectral methods by substantial simulated networks and eight real-world networks.
Community Detection by Principal Components Clustering Methods
Based on the classical Degree Corrected Stochastic Blockmodel (DCSBM) model for network community detection problem, we propose two novel approaches: principal component clustering (PCC) and normalized principal component clustering (NPCC). Without any parameters to be estimated, the PCC method is simple to be implemented. Under mild conditions, we show that PCC yields consistent community detection. NPCC is designed based on the combination of the PCC and the RSC method (Qin & Rohe 2013). Population analysis for NPCC shows that NPCC returns perfect clustering for the ideal case under DCSBM. PCC and NPCC is illustrated through synthetic and real-world datasets. Numerical results show that NPCC provides a significant improvement compare with PCC and RSC. Moreover, NPCC inherits nice properties of PCC and RSC such that NPCC is insensitive to the number of eigenvectors to be clustered and the choosing of the tuning parameter. When dealing with two weak signal networks Simmons and Caltech, by considering one more eigenvectors for clustering, we provide two refinements PCC+ and NPCC+ of PCC and NPCC, respectively. Both two refinements algorithms provide improvement performances compared with their original algorithms. Especially, NPCC+ provides satisfactory performances on Simmons and Caltech, with error rates of 121/1137 and 96/590, respectively.
Mitigating Bias in Set Selection with Noisy Protected Attributes
Mehrotra, Anay, Celis, L. Elisa
Subset selection algorithms are ubiquitous in AI-driven applications, including, online recruiting portals and image search engines, so it is imperative that these tools are not discriminatory on the basis of protected attributes such as gender or race. Currently, fair subset selection algorithms assume that the protected attributes are known as part of the dataset. However, attributes may be noisy due to errors during data collection or if they are imputed (as is often the case in real-world settings). While a wide body of work addresses the effect of noise on the performance of machine learning algorithms, its effect on fairness remains largely unexamined. We find that in the presence of noisy protected attributes, in attempting to increase fairness without considering noise, one can, in fact, decrease the fairness of the result! Towards addressing this, we consider an existing noise model in which there is probabilistic information about the protected attributes (e.g.,[19, 32, 56, 44]), and ask is fair selection is possible under noisy conditions? We formulate a ``denoised'' selection problem which functions for a large class of fairness metrics; given the desired fairness goal, the solution to the denoised problem violates the goal by at most a small multiplicative amount with high probability. Although the denoised problem turns out to be NP-hard, we give a linear-programming based approximation algorithm for it. We empirically evaluate our approach on both synthetic and real-world datasets. Our empirical results show that this approach can produce subsets which significantly improve the fairness metrics despite the presence of noisy protected attributes, and, compared to prior noise-oblivious approaches, has better Pareto-tradeoffs between utility and fairness.
Classifier-independent Lower-Bounds for Adversarial Robustness
We theoretically analyse the limits of robustness to test-time adversarial and noisy examples in classification. Our work focuses on deriving bounds which uniformly apply to all classifiers (i.e all measurable functions from features to labels) for a given problem. Our contributions are two-fold. (1) We use optimal transport theory to derive variational formulae for the Bayes-optimal error a classifier can make on a given classification problem, subject to adversarial attacks. The optimal adversarial attack is then an optimal transport plan for a certain binary cost-function induced by the specific attack model, and can be computed via a simple algorithm based on maximal matching on bipartite graphs. (2) We derive explicit lower-bounds on the Bayes-optimal error in the case of the popular distance-based attacks. These bounds are universal in the sense that they depend on the geometry of the class-conditional distributions of the data, but not on a particular classifier. Our results are in sharp contrast with the existing literature, wherein adversarial vulnerability of classifiers is derived as a consequence of nonzero ordinary test error.
Distance-Based Anomaly Detection for Industrial Surfaces Using Triplet Networks
Tayeh, Tareq, Aburakhia, Sulaiman, Myers, Ryan, Shami, Abdallah
Surface anomaly detection plays an important quality control role in many manufacturing industries to reduce scrap production. Machine-based visual inspections have been utilized in recent years to conduct this task instead of human experts. In particular, deep learning Convolutional Neural Networks (CNNs) have been at the forefront of these image processing-based solutions due to their predictive accuracy and efficiency. Training a CNN on a classification objective requires a sufficiently large amount of defective data, which is often not available. In this paper, we address that challenge by training the CNN on surface texture patches with a distance-based anomaly detection objective instead. A deep residual-based triplet network model is utilized, and defective training samples are synthesized exclusively from non-defective samples via random erasing techniques to directly learn a similarity metric between the same-class samples and out-of-class samples. Evaluation results demonstrate the approach's strength in detecting different types of anomalies, such as bent, broken, or cracked surfaces, for known surfaces that are part of the training data and unseen novel surfaces.
An Experimentation Platform for Explainable Coalition Situational Understanding
Barrett-Powell, Katie, Furby, Jack, Hiley, Liam, Vilamala, Marc Roig, Taylor, Harrison, Cerutti, Federico, Preece, Alun, Xing, Tianwei, Garcia, Luis, Srivastava, Mani, Braines, Dave
Therefore, our work alliances through multiple means: diplomatic, economic, seeks to advance capabilities in explainable AI/ML to allow conventional and unconventional warfare, including information a human operative to'calibrate their trust' in an AI/ML asset warfare. A critical requirement for allies is potentially provided by a different coalition partner (Tomsett rapid and continuous integration of capabilities to collect, et al. 2020). The purpose of human-machine teaming is process, disseminate and exploit actionable information and to aim for each party to exploit the strengths of, and compensate intelligence. To achieve this, the MDO layered ISR concept for the weaknesses of, the other (Cummings 2014).
On Regulating AI in Medical Products (OnRAMP)
Medical AI products require certification before deployment in most jurisdictions. To date, no clear pathways for regulating medical AI exist. I present a methodological guide to the development of a regulatory package which will form part of a certification process. This approach is predicated on the translation between a statistical risk perspective, typical of medical device regulators, and a deep understanding of machine learning methodologies. This work of translation envisages the statistician as the key negotiator between medical device regulators and machine learning experts, allowing them to communicate more clearly, and thus lead to the development of standardised pathways for medical AI regulation.