Government
Knowledge-Integrated Informed AI for National Security
The state of artificial intelligence technology has a rich history that dates back decades and includes two fall-outs before the explosive resurgence of today, which is credited largely to data-driven techniques. While AI technology has and continues to become increasingly mainstream with impact across domains and industries, it's not without several drawbacks, weaknesses, and potential to cause undesired effects. AI techniques are numerous with many approaches and variants, but they can be classified simply based on the degree of knowledge they capture and how much data they require; two broad categories emerge as prominent across AI to date: (1) techniques that are primarily, and often solely, data-driven while leveraging little to no knowledge and (2) techniques that primarily leverage knowledge and depend less on data. Now, a third category is starting to emerge that leverages both data and knowledge, that some refer to as "informed AI." This third category can be a game changer within the national security domain where there is ample scientific and domain-specific knowledge that stands ready to be leveraged, and where purely data-driven AI can lead to serious unwanted consequences.
FedQAS: Privacy-aware machine reading comprehension with federated learning
Ait-Mlouk, Addi, Alawadi, Sadi, Toor, Salman, Hellander, Andreas
Machine reading comprehension (MRC) of text data is one important task in Natural Language Understanding. It is a complex NLP problem with a lot of ongoing research fueled by the release of the Stanford Question Answering Dataset (SQuAD) and Conversational Question Answering (CoQA). It is considered to be an effort to teach computers how to "understand" a text, and then to be able to answer questions about it using deep learning. However, until now large-scale training on private text data and knowledge sharing has been missing for this NLP task. Hence, we present FedQAS, a privacy-preserving machine reading system capable of leveraging large-scale private data without the need to pool those datasets in a central location. The proposed approach combines transformer models and federated learning technologies. The system is developed using the FEDn framework and deployed as a proof-of-concept alliance initiative. FedQAS is flexible, language-agnostic, and allows intuitive participation and execution of local model training. In addition, we present the architecture and implementation of the system, as well as provide a reference evaluation based on the SQUAD dataset, to showcase how it overcomes data privacy issues and enables knowledge sharing between alliance members in a Federated learning setting.
A Coupled CP Decomposition for Principal Components Analysis of Symmetric Networks
Weylandt, Michael, Michailidis, George
In a number of application domains, one observes a sequence of network data; for example, repeated measurements between users interactions in social media platforms, financial correlation networks over time, or across subjects, as in multi-subject studies of brain connectivity. One way to analyze such data is by stacking networks into a third-order array or tensor. We propose a principal components analysis (PCA) framework for sequence network data, based on a novel decomposition for semi-symmetric tensors. We derive efficient algorithms for computing our proposed "Coupled CP" decomposition and establish estimation consistency of our approach under an analogue of the spiked covariance model with rates the same as the matrix case up to a logarithmic term. Our framework inherits many of the strengths of classical PCA and is suitable for a wide range of unsupervised learning tasks, including identifying principal networks, isolating meaningful changepoints or outliers across observations, and for characterizing the "variability network" of the most varying edges. Finally, we demonstrate the effectiveness of our proposal on simulated data and on examples from political science and financial economics. The proof techniques used to establish our main consistency results are surprisingly straight-forward and may find use in a variety of other matrix and tensor decomposition problems.
Predictive Inference with Weak Supervision
Cauchois, Maxime, Gupta, Suyash, Ali, Alnur, Duchi, John
Consider the typical supervised learning pipeline that we teach students learning statistical machine learning: we collect data in (X, Y) pairs, where Y is a label or target to be predicted; we pick a model and loss measuring the fidelity of the model to observed data; we choose the model minimizing the loss and validate it on held-out data. This picture obscures what is becoming one of the major challenges in this endeavor: that of actually collecting highquality labeled data [44, 13, 38]. Hand labeling large-scale training sets is often impractically expensive. Consider, as simple motivation, a ranking problem: a prediction is an ordered list of a set of items, yet available feedback is likely to be incomplete and partial, such as a top element (for example, in web search a user clicks on a single preferred link, or in a grocery, an individual buys one kind of milk but provides no feedback on the other brands present). Developing methods to leverage such partial and weak feedback is therefore becoming a major focus, and researchers have developed methods to transform weak and noisy labels into a dataset with strong, "gold-standard" labels [38, 56]. In this paper, we adopt this weakly labeled setting, but instead of considering model fitting and the construction of strong labels, we focus on validation, model confidence, and predictive inference, moving beyond point predictions and single labels. Our goal is to develop methods to rigorously quantify the confidence a practitioner should have in a model given only weak labels.
What Did the IRS Want With Your Selfies?
In November, the IRS signed an $86 million contract with identity verification startup ID.me, announcing that it would require taxpayers to provide personal, identifying materials, including selfies, to access their tax records. Privacy and civil rights advocates responded immediately, forming a coalition of close to 20 groups--from the National Lawyers Guild to the Council on American-Islamic Relations--that criticized the "destructive results of facial recognition technologyโฆfrom police using it to track Black Lives Matter protesters, to wrongful arrests, to manipulative marketing." The IRS plan, those groups said, "would have expanded the scope of these harms and impact the lives of millions more people."
The IRS Drops Facial Recognition Verification After Uproar
The Internal Revenue Service is dropping a controversial facial recognition system that requires people to upload video selfies when creating new IRS online accounts. This story originally appeared on Ars Technica, a trusted source for technology news, tech policy analysis, reviews, and more. Ars is owned by WIRED's parent company, Condรฉ Nast. "The IRS announced it will transition away from using a third-party service for facial recognition to help authenticate people creating new online accounts," the agency said on Monday. "The transition will occur over the coming weeks in order to prevent larger disruptions to taxpayers during filing season. During the transition, the IRS will quickly develop and bring online an additional authentication process that does not involve facial recognition."
Microsoft and Sony are buying up the video game world. The FTC could stop them.
On the same day the two gaming goliaths announced the deal, the FTC and DOJ launched a joint public inquiry with the goal of better detecting and preventing anti-competitive deals. Shortly after, Bloomberg reported the FTC had assumed responsibility for reviewing the Microsoft and Activision Blizzard deal. The FTC declined to comment or confirm an existing investigation. Stoller, though, pointed out that the stock market appears to be reacting to the looming specter of this investigation. "Activision is [trading at] $80, and the purchase price is at $95," he said.
Chrpa
Effective and efficient reasoning in adversarial environments is important for many real-world applications ranging from cybersecurity to military operations. Deliberative reasoning techniques, such as Automated Planning, often restrict to static environments where only an agent can make changes by its actions. On the other hand, such techniques are effective and can generate non-trivial solutions. To explicitly reason in environments with an active adversary such as zero-sum games, the game-theoretic framework such as the Double Oracle algorithm can be leveraged. In this paper, we leverage the notions of critical and adversary actions, where critical actions should be applied before the adversary ones. We propose heuristics that provide a guidance for planners about what (critical) actions and in which order have to be applied in a good plan.
Hampton
Navigating a career constitutes one of life's most enduring challenges, particularly within a unique organization like the US Navy. While the Navy has numerous resources for guidance, accessing and identifying key information sources across the many existing platforms can be challenging for sailors (e.g., determining the appropriate program or point of contact, developing an accurate understanding of the process, and even recognizing the need for planning itself). Focusing on intermediate goals, evaluations, education, certifications, and training is quite demanding, even before considering their cumulative long-term implications. These are on top of generic personal issues, such as financial difficulties and homesickness when at sea for prolonged periods. We present the preliminary construction of a conversational intelligent agent designed to provide a user-friendly, adaptive environment that recognizes user input pertinent to these issues and provides guidance to appropriate resources within the Navy.
Machine learning fine-tunes graphene synthesis
Rice University chemists are employing machine learning to fine-tune its flash Joule heating process to make graphene. A flash signifies the creation of graphene from waste. Rice University scientists are using machine learning techniques to streamline the process of synthesizing graphene from waste through flash Joule heating. This flash Joule process has expanded beyond making graphene from various carbon sources, to extracting other materials, like metals, from urban waste. The technique is the same for all of the above: blasting a jolt of high energy through the source material to eliminate all but the desired product.