Government
Fast Distributionally Robust Learning with Variance Reduced Min-Max Optimization
Yu, Yaodong, Lin, Tianyi, Mazumdar, Eric, Jordan, Michael I.
With machine learning systems increasingly being deployed in real-world settings, there is an urgent need for machine learning approaches that can adapt to--or are robust to--changes in the environment. Despite this, the dominant paradigm for supervised learning [Hastie et al., 2009] remains that of empirical risk minimization (ERM) [Vapnik, 2013], wherein a model is trained by minimizing a loss over a fixed set of training data. A key assumption underlying this approach is that the training data is from the same distribution as the test data--i.e., the distribution of the data does not change between training time and deployment. Such assumptions are well known to rarely hold in practice. Indeed, the distribution may change due to sample selection bias, nonstationarity in the environment [Quionero-Candela et al., 2009], or even adversarial perturbations [Szegedy et al., 2013, Madry et al., 2018], leaving machine learning models trained through ERM particularly susceptible to adversarial attacks [Szegedy et al., 2013, Carlini and Wagner, 2017] or to degraded performance from distribution shifts. Distributionally robust supervised learning (DRSL) seeks to address this issue by explicitly optimizing for solutions that are are robust to adversarial distribution shifts.
Detection of Signal in the Spiked Rectangular Models
Jung, Ji Hyung, Chung, Hye Won, Lee, Ji Oon
We consider the problem of detecting signals in the rank-one signal-plus-noise data matrix models that generalize the spiked Wishart matrices. We show that the principal component analysis can be improved by pre-transforming the matrix entries if the noise is non-Gaussian. As an intermediate step, we prove a sharp phase transition of the largest eigenvalues of spiked rectangular matrices, which extends the Baik-Ben Arous-P\'ech\'e (BBP) transition. We also propose a hypothesis test to detect the presence of signal with low computational complexity, based on the linear spectral statistics, which minimizes the sum of the Type-I and Type-II errors when the noise is Gaussian.
TRECVID 2020: A comprehensive campaign for evaluating video retrieval tasks across multiple application domains
Awad, George, Butt, Asad A., Curtis, Keith, Fiscus, Jonathan, Godil, Afzal, Lee, Yooyoung, Delgado, Andrew, Zhang, Jesse, Godard, Eliot, Chocot, Baptiste, Diduch, Lukas, Liu, Jeffrey, Smeaton, Alan F., Graham, Yvette, Jones, Gareth J. F., Kraaij, Wessel, Quenot, Georges
The TREC Video Retrieval Evaluation (TRECVID) is a TREC-style video analysis and retrieval evaluation with the goal of promoting progress in research and development of content-based exploitation and retrieval of information from digital video via open, metrics-based evaluation. Over the last twenty years this effort has yielded a better understanding of how systems can effectively accomplish such processing and how one can reliably benchmark their performance. TRECVID has been funded by NIST (National Institute of Standards and Technology) and other US government agencies. In addition, many organizations and individuals worldwide contribute significant time and effort. TRECVID 2020 represented a continuation of four tasks and the addition of two new tasks. In total, 29 teams from various research organizations worldwide completed one or more of the following six tasks: 1. Ad-hoc Video Search (AVS), 2. Instance Search (INS), 3. Disaster Scene Description and Indexing (DSDI), 4. Video to Text Description (VTT), 5. Activities in Extended Video (ActEV), 6. Video Summarization (VSUM). This paper is an introduction to the evaluation framework, tasks, data, and measures used in the evaluation campaign.
Proceedings - AI/ML for Cybersecurity: Challenges, Solutions, and Novel Ideas at SIAM Data Mining 2021
Emanuello, John, Ferguson-Walter, Kimberly, Hemberg, Erik, Reilly, Una-May O, Ridley, Ahmad, Ross, Dennis, Staheli, Diane, Streilein, William
Malicious cyber activity is ubiquitous and its harmful effects have dramatic and often irreversible impacts on society. Given the shortage of cybersecurity professionals, the ever-evolving adversary, the massive amounts of data which could contain evidence of an attack, and the speed at which defensive actions must be taken, innovations which enable autonomy in cybersecurity must continue to expand, in order to move away from a reactive defense posture and towards a more proactive one. The challenges in this space are quite different from those associated with applying AI in other domains such as computer vision. The environment suffers from an incredibly high degree of uncertainty, stemming from the intractability of ingesting all the available data, as well as the possibility that malicious actors are manipulating the data. Another unique challenge in this space is the dynamism of the adversary causes the indicators of compromise to change frequently and without warning. In spite of these challenges, machine learning has been applied to this domain and has achieved some success in the realm of detection. While this aspect of the problem is far from solved, a growing part of the commercial sector is providing ML-enhanced capabilities as a service. Many of these entities also provide platforms which facilitate the deployment of these automated solutions. Academic research in this space is growing and continues to influence current solutions, as well as strengthen foundational knowledge which will make autonomous agents in this space a possibility.
Random Projections for Improved Adversarial Robustness
Carbone, Ginevra, Sanguinetti, Guido, Bortolussi, Luca
We propose two training techniques for improving the robustness of Neural Networks to adversarial attacks, i.e. manipulations of the inputs that are maliciously crafted to fool networks into incorrect predictions. Both methods are independent of the chosen attack and leverage random projections of the original inputs, with the purpose of exploiting both dimensionality reduction and some characteristic geometrical properties of adversarial perturbations. The first technique is called RP-Ensemble and consists of an ensemble of networks trained on multiple projected versions of the original inputs. The second one, named RP-Regularizer, adds instead a regularization term to the training objective.
Tesla CEO Elon Musk blasts reports blaming Autopilot for deadly Model S crash as 'completely false'
Tesla executives defended the automaker's semi-self-driving system on Monday after it came under scrutiny following a deadly crash involving a Tesla Model S in Texas this month. CEO Elon Musk rejected suggestions that the company's Autopilot was to blame. "This is completely false," he said, adding that journalists who suggested Autopilot was at fault "should be ashamed of themselves." After the crash, Harris County Precinct 4 Constable Mark Herman told multiple outlets, including the Wall Street Journal and Consumer Reports, that investigators were 99.9% sure that no one was behind the wheel when the vehicle crashed. The National Transportation Safety Board and the National Highway Traffic Safety Administration are investigating the crash.
Stunning DDT dump site off L.A. coast much bigger than scientists expected
When the research vessel Sally Ride set sail for Santa Catalina Island to map an underwater graveyard of DDT waste barrels, its crew had high hopes of documenting for the first time just how many corroded containers littered the seafloor off the coast of Los Angeles. But as the scientists on deck began interpreting sonar images gathered by two deep-sea robots, they were quickly overwhelmed. It was like trying to count stars in the Milky Way. The dumpsite it turned out, was much, much bigger than expected. After spending two weeks surveying a swath of seafloor larger than the city of San Francisco, the scientists could find no end to the dumping ground.
AI 50: America's Most Promising Artificial Intelligence Companies
The Covid-19 pandemic was devastating for many industries, but it only accelerated the use of artificial intelligence across the U.S. economy. Amid the crisis, companies scrambled to create new services for remote workers and students, beef up online shopping and dining options, make customer call centers more efficient and speed development of important new drugs. Even as applications of machine learning and perception platforms become commonplace, a thick layer of hype and fuzzy jargon clings to AI-enabled software.That makes it tough to identify the most compelling companies in the space--especially those finding new ways to use AI that create value by making humans more efficient, not redundant. With this in mind, Forbes has partnered with venture firms Sequoia Capital and Meritech Capital to create our third annual AI 50, a list of private, promising North American companies that are using artificial intelligence in ways that are fundamental to their operations. To be considered, businesses must be privately-held and utilizing machine learning (where systems learn from data to improve on tasks), natural language processing (which enables programs to "understand" written or spoken language) or computer vision (which relates to how machines "see"). AI companies incubated at, largely funded through or acquired by large tech, manufacturing or industrial firms aren't eligible for consideration. Our list was compiled through a submission process open to any AI company in the U.S. and Canada. The application asked companies to provide details on their technology, business model, customers and financials like funding, valuation and revenue history (companies had the option to submit information confidentially, to encourage greater transparency). Forbes received several hundred entries, of which nearly 400 qualified for consideration. From there, our data partners applied an algorithm to identify 100 companies with the highest quantitative scores--and that also made diversity a priority. Next, a panel of expert AI judges evaluated the finalists to find the 50 most compelling companies (they were precluded from judging companies in which they have a vested interest). Among trends this year are what Sequoia Capital's Konstantine Buhler calls AI workbench companies--building of platforms tailored to different enterprises, including Dataiku, DataRobot Domino Data and Databricks.
AI- and ML-enabled spectrum management tech goal of DoD research - Military Embedded Systems
The U.S. Department of Defense (DoD) has issued a third Request for Prototype Proposal (RPPs) in support of electromagnetic spectrum research related to the capabilities of the over 400 members of the National Spectrum Consortium. Issued under the DoD's Spectrum Access Research & Development (SAR&DP) Program, the RPP is part of a series of requirements to develop near real time spectrum management technologies that leverage machine learning (ML) and artificial intelligence (AI) to more efficiently allocate spectrum assignments based on operational planning and intended operational outcomes. Officials claim that this specific RPP is centered on the Operational Spectrum Comprehension, Analytics, and Response (OSCAR) effort. This project will aim to create a software application with unified graphical user interface, automated workflows, sensor network, and extensible framework needed at testing and training ranges for aerial combat training to ensure that spectrum is available when and where needed for AWS-3 impacted systems and incumbent systems. According to officials, the goal is to provide advanced spectrum management capabilities to the incumbent systems in the AWS-3 bands; however, this prototype will be applicable to all spectrum being managed on range.
The EU Is Proposing Regulations On AI--And The Impact On Healthcare Could Be Significant
The emphasis and development of artificial intelligence (AI) is swiftly growing, with innovators across the globe trying to create more viable use-cases for this groundbreaking technology. AI's market reach has penetrated nearly every large industry, including manufacturing, retail, infrastructure, financial services, defense, and healthcare, among countless other sectors. Healthcare especially has experienced an incredible amount of attention in the AI space. The value proposition of AI in healthcare is undoubtedly extensive, especially as the industry is poised to surpass over $11 trillion in market valuation, and given that healthcare is such an inherently data-rich, innovation heavy, and operationally nuanced field. Last week, the European Union (EU) put forth its "Proposal for a Regulation on a European approach for Artificial Intelligence," intending to create "the first ever legal framework on AI, which addresses the risks of AI and positions Europe to play a leading role globally."