Government
The Mail
Adam Gopnik, in his essay about the legacy of the Marquis de Lafayette, points out that Americans rarely understand why Lafayette does not enjoy the exalted reputation in France that he does in America (Books, August 23rd). As Gopnik mentions, supporters of the French Revolution blamed Lafayette for not preventing the royal family's flight from France, in June, 1791. He was responsible, as the commander of the Paris militia, for security at the palace. When the King's disappearance became known, Lafayette colluded in saying that he had been kidnapped, an explanation that quickly fell apart when the King's denunciation of the Revolution was published. Soon afterward, protesters gathered to rally behind a petition objecting to the restoration of the King to the throne, and the National Guard was ordered to disperse them.
Why AI and Automation Provide Superhuman Security
Every CISO's worst nightmare is that their organization will become the victim of a cyberattack. Unfortunately, this is a scenario that is becoming increasingly likely every day, as threats actors are ready to exploit new and more sophisticated vectors. For example, supply-chain-based attacks, such as the SolarWinds SUNBURST attack, are not simple system vulnerabilities. These attacks consist of a complex series of actions in which the initial infection is just the first step and are both more and more commonplace and difficult to prevent. This particular attack swept the globe since the campaign activation in March 2020, targeting the finance, government, healthcare, education and infrastructure verticals alike.
Sending drones thousands of miles away without proper intelligence isn't good enough: Rep. Waltz
Rep. Mike Waltz and former national security adviser to VP Pence Gen. Keith Kellogg react to the mistaken drone strike killing innocent civilians. "Sunday Night in America" host Trey Gowdy discussed the recent reveal that the drone strike in Kabul touted by the Biden administration struck down civilians including seven children. "When you're looking at a drone strike, those drones could basically read a license plate on a car, that's how good their optics are," Kellogg said. "My concerns are that just a couple of days after the ISIS-K strike that killed 13 great Americans and wounded 20, I think there was a big push to try to get these planners. Remember, we were told these were ISIS-K planners when they weren't. They were tracking a white van for approximately 8 hours, then they said it was a righteous kill. And they were dead wrong about it."
Molecular Energy Learning Using Alternative Blackbox Matrix-Matrix Multiplication Algorithm for Exact Gaussian Process
Sun, Jiace, Cheng, Lixue, Miller, Thomas F. III
We present an application of the blackbox matrix-matrix multiplication (BBMM) algorithm to scale up the Gaussian Process (GP) training of molecular energies in the molecular-orbital based machine learning (MOB-ML) framework. An alternative implementation of BBMM (AltBBMM) is also proposed to train more efficiently (over four-fold speedup) with the same accuracy and transferability as the original BBMM implementation. The training of MOB-ML was limited to 220 molecules, and BBMM and AltBBMM scale the training of MOB-ML up by over 30 times to 6500 molecules (more than a million pair energies). The accuracy and transferability of both algorithms are examined on the benchmark datasets of organic molecules with 7 and 13 heavy atoms. These lower-scaling implementations of the GP preserve the state-of-the-art learning efficiency in the low-data regime while extending it to the large-data regime with better accuracy than other available machine learning works on molecular energies.
The Case for Claim Difficulty Assessment in Automatic Fact Checking
Singh, Prakhar, Das, Anubrata, Li, Junyi Jessy, Lease, Matthew
Fact-checking is the process (human, automated, or hybrid) by which claims (i.e., purported facts) are evaluated for veracity. In this article, we raise an issue that has received little attention in prior work - that some claims are far more difficult to fact-check than others. We discuss the implications this has for both practical fact-checking and research on automated fact-checking, including task formulation and dataset design. We report a manual analysis undertaken to explore factors underlying varying claim difficulty and categorize several distinct types of difficulty. We argue that prediction of claim difficulty is a missing component of today's automated fact-checking architectures, and we describe how this difficulty prediction task might be split into a set of distinct subtasks.
"Hello, It's Me": Deep Learning-based Speech Synthesis Attacks in the Real World
Wenger, Emily, Bronckers, Max, Cianfarani, Christian, Cryan, Jenna, Sha, Angela, Zheng, Haitao, Zhao, Ben Y.
Advances in deep learning have introduced a new wave of voice synthesis tools, capable of producing audio that sounds as if spoken by a target speaker. If successful, such tools in the wrong hands will enable a range of powerful attacks against both humans and software systems (aka machines). This paper documents efforts and findings from a comprehensive experimental study on the impact of deep-learning based speech synthesis attacks on both human listeners and machines such as speaker recognition and voice-signin systems. We find that both humans and machines can be reliably fooled by synthetic speech and that existing defenses against synthesized speech fall short. These findings highlight the need to raise awareness and develop new protections against synthetic speech for both humans and machines.
Assessing the quality of sources in Wikidata across languages: a hybrid approach
Amaral, Gabriel, Piscopo, Alessandro, Kaffee, Lucie-Aimée, Rodrigues, Odinaldo, Simperl, Elena
Wikidata is one of the most important sources of structured data on the web, built by a worldwide community of volunteers. As a secondary source, its contents must be backed by credible references; this is particularly important as Wikidata explicitly encourages editors to add claims for which there is no broad consensus, as long as they are corroborated by references. Nevertheless, despite this essential link between content and references, Wikidata's ability to systematically assess and assure the quality of its references remains limited. To this end, we carry out a mixed-methods study to determine the relevance, ease of access, and authoritativeness of Wikidata references, at scale and in different languages, using online crowdsourcing, descriptive statistics, and machine learning. Building on previous work of ours, we run a series of microtasks experiments to evaluate a large corpus of references, sampled from Wikidata triples with labels in several languages. We use a consolidated, curated version of the crowdsourced assessments to train several machine learning models to scale up the analysis to the whole of Wikidata. The findings help us ascertain the quality of references in Wikidata, and identify common challenges in defining and capturing the quality of user-generated multilingual structured data on the web. We also discuss ongoing editorial practices, which could encourage the use of higher-quality references in a more immediate way. All data and code used in the study are available on GitHub for feedback and further improvement and deployment by the research community.
Language Models as a Knowledge Source for Cognitive Agents
Wray,, Robert E. III, Kirk, James R., Laird, John E.
Language models (LMs) are sentence-completion engines trained on massive corpora. LMs have emerged as a significant breakthrough in natural-language processing, providing capabilities that go far beyond sentence completion including question answering, summarization, and natural-language inference. While many of these capabilities have potential application to cognitive systems, exploiting language models as a source of task knowledge, especially for task learning, offers significant, near-term benefits. We introduce language models and the various tasks to which they have been applied and then review methods of knowledge extraction from language models. The resulting analysis outlines both the challenges and opportunities for using language models as a new knowledge source for cognitive systems. It also identifies possible ways to improve knowledge extraction from language models using the capabilities provided by cognitive systems. Central to success will be the ability of a cognitive agent to itself learn an abstract model of the knowledge implicit in the LM as well as methods to extract high-quality knowledge effectively and efficiently. To illustrate, we introduce a hypothetical robot agent and describe how language models could extend its task knowledge and improve its performance and the kinds of knowledge and methods the agent can use to exploit the knowledge within a language model.
Sharp global convergence guarantees for iterative nonconvex optimization: A Gaussian process perspective
Chandrasekher, Kabir Aladin, Pananjady, Ashwin, Thrampoulidis, Christos
We consider a general class of regression models with normally distributed covariates, and the associated nonconvex problem of fitting these models from data. We develop a general recipe for analyzing the convergence of iterative algorithms for this task from a random initialization. In particular, provided each iteration can be written as the solution to a convex optimization problem satisfying some natural conditions, we leverage Gaussian comparison theorems to derive a deterministic sequence that provides sharp upper and lower bounds on the error of the algorithm with sample-splitting. Crucially, this deterministic sequence accurately captures both the convergence rate of the algorithm and the eventual error floor in the finite-sample regime, and is distinct from the commonly used "population" sequence that results from taking the infinite-sample limit. We apply our general framework to derive several concrete consequences for parameter estimation in popular statistical models including phase retrieval and mixtures of regressions. Provided the sample size scales near-linearly in the dimension, we show sharp global convergence rates for both higher-order algorithms based on alternating updates and first-order algorithms based on subgradient descent. These corollaries, in turn, yield multiple consequences, including: (a) Proof that higher-order algorithms can converge significantly faster than their first-order counterparts (and sometimes super-linearly), even if the two share the same population update and (b) Intricacies in super-linear convergence behavior for higher-order algorithms, which can be nonstandard (e.g., with exponent 3/2) and sensitive to the noise level in the problem. We complement these results with extensive numerical experiments, which show excellent agreement with our theoretical predictions.
Reconstructing Cosmic Polarization Rotation with ResUNet-CMB
Cosmic polarization rotation, which may result from parity-violating new physics or the presence of primordial magnetic fields, converts $E$-mode polarization of the cosmic microwave background (CMB) into $B$-mode polarization. Anisotropic cosmic polarization rotation leads to statistical anisotropy in CMB polarization and can be reconstructed with quadratic estimator techniques similar to those designed for gravitational lensing of the CMB. At the sensitivity of upcoming CMB surveys, lensing-induced $B$-mode polarization will act as a limiting factor in the search for anisotropic cosmic polarization rotation, meaning that an analysis which incorporates some form of delensing will be required to improve constraints on the effect with future surveys. In this paper we extend the ResUNet-CMB convolutional neural network to reconstruct anisotropic cosmic polarization rotation in the presence of gravitational lensing and patchy reionization, and we show that the network simultaneously reconstructs all three effects with variance that is lower than that from the standard quadratic estimator nearly matching the performance of an iterative reconstruction method.