Government
Using the power of Machine Learning to detect cyber attacks - Express Computer
As the world becomes increasingly digital, we are unlocking more value and growth than ever before. However, a challenge that governments, enterprises and well as individuals leveraging technology are constantly facing is the growing threat of cyberattacks that looms large over us. Cyber security solutions provider SonicWall's 2019 report revealed 10.52 billion malware attacks in 2018, a 217% increase in IoT attacks and 391,689 new variants of attack that were identified. What's more is that cyber criminals today are evolving with technology and upping their game. Such incidents don't just have the potential to bring businesses to a standstill but can also inflict serious damages to their resources and repute.
Assessing regulatory fairness through machine learning
The analysis, published this week in the proceedings of the Association of Computing Machinery Conference on Fairness, Accountability and Transparency(link is external), evaluates machine learning techniques designed to support a U.S. Environmental Protection Agency (EPA) initiative to reduce severe violations of the Clean Water Act. It reveals how two key elements of so-called algorithmic design influence which communities are targeted for compliance efforts and, consequently, who bears the burden of pollution violations. The analysis -- funded through the Stanford Woods Institute for the Environment's Realizing Environmental Innovation Program -- is timely given recent executive actions(link is external) calling for renewed focus on environmental justice. "Machine learning is being used to help manage an overwhelming number of things that federal agencies are tasked to do -- as a way to help increase efficiency," said study co-principal investigator Daniel Ho, the William Benjamin Scott and Luna M. Scott Professor of Law at Stanford Law School. "Yet what we also show is that simply designing a machine learning-based system can have an additional benefit."
Interpretable Data-driven Methods for Subgrid-scale Closure in LES for Transcritical LOX/GCH4 Combustion
Chung, Wai Tong, Mishra, Aashwin Ananda, Ihme, Matthias
Many practical combustion systems such as those in rockets, gas turbines, and internal combustion engines operate under high pressures that surpass the thermodynamic critical limit of fuel-oxidizer mixtures. These conditions require the consideration of complex fluid behaviors that pose challenges for numerical simulations, casting doubts on the validity of existing subgrid-scale (SGS) models in large-eddy simulations of these systems. While data-driven methods have shown high accuracy as closure models in simulations of turbulent flames, these models are often criticized for lack of physical interpretability, wherein they provide answers but no insight into their underlying rationale. The objective of this study is to assess SGS stress models from conventional physics-driven approaches and an interpretable machine learning algorithm, i.e., the random forest regressor, in a turbulent transcritical non-premixed flame. To this end, direct numerical simulations (DNS) of transcritical liquid-oxygen/gaseous-methane (LOX/GCH4) inert and reacting flows are performed. Using this data, a priori analysis is performed on the Favre-filtered DNS data to examine the accuracy of physics-based and random forest SGS-models under these conditions. SGS stresses calculated with the gradient model show good agreement with the exact terms extracted from filtered DNS. The accuracy of the random-forest regressor decreased when physics-based constraints are applied to the feature set. Results demonstrate that random forests can perform as effectively as algebraic models when modeling subgrid stresses, only when trained on a sufficiently representative database. The employment of random forest feature importance score is shown to provide insight into discovering subgrid-scale stresses through sparse regression.
Multi-Class Multiple Instance Learning for Predicting Precursors to Aviation Safety Events
Bleu-Laine, Marc-Henri, Puranik, Tejas G., Mavris, Dimitri N., Matthews, Bryan
In recent years, there has been a rapid growth in the application of machine learning techniques that leverage aviation data collected from commercial airline operations to improve safety. Anomaly detection and predictive maintenance have been the main targets for machine learning applications. However, this paper focuses on the identification of precursors, which is a relatively newer application. Precursors are events correlated with adverse events that happen prior to the adverse event itself. Therefore, precursor mining provides many benefits including understanding the reasons behind a safety incident and the ability to identify signatures, which can be tracked throughout a flight to alert the operators of the potential for an adverse event in the future. This work proposes using the multiple-instance learning (MIL) framework, a weakly supervised learning task, combined with carefully designed binary classifier leveraging a Multi-Head Convolutional Neural Network-Recurrent Neural Network (MHCNN-RNN) architecture. Multi-class classifiers are then created and compared, enabling the prediction of different adverse events for any given flight by combining binary classifiers, and by modifying the MHCNN-RNN to handle multiple outputs. Results obtained showed that the multiple binary classifiers perform better and are able to accurately forecast high speed and high path angle events during the approach phase. Multiple binary classifiers are also capable of determining the aircraft's parameters that are correlated to these events. The identified parameters can be considered precursors to the events and may be studied/tracked further to prevent these events in the future.
Empirical Mode Modeling: A data-driven approach to recover and forecast nonlinear dynamics from noisy data
Park, Joseph, Pao, Gerald M, Stabenau, Erik, Sugihara, George, Lorimer, Thomas
Data-driven, model-free analytics are natural choices for discovery and forecasting of complex, nonlinear systems. Methods that operate in the system state-space require either an explicit multidimensional state-space, or, one approximated from available observations. Since observational data are frequently sampled with noise, it is possible that noise can corrupt the state-space representation degrading analytical performance. Here, we evaluate the synthesis of empirical mode decomposition with empirical dynamic modeling, which we term empirical mode modeling, to increase the information content of state-space representations in the presence of noise. Evaluation of a mathematical, and, an ecologically important geophysical application across three different state-space representations suggests that empirical mode modeling may be a useful technique for data-driven, model-free, state-space analysis in the presence of noise.
Topical Language Generation using Transformers
Zandie, Rohola, Mahoor, Mohammad H.
Large-scale transformer-based language models (LMs) demonstrate impressive capabilities in open text generation. However, controlling the generated text's properties such as the topic, style, and sentiment is challenging and often requires significant changes to the model architecture or retraining and fine-tuning the model on new supervised data. This paper presents a novel approach for Topical Language Generation (TLG) by combining a pre-trained LM with topic modeling information. We cast the problem using Bayesian probability formulation with topic probabilities as a prior, LM probabilities as the likelihood, and topical language generation probability as the posterior. In learning the model, we derive the topic probability distribution from the user-provided document's natural structure. Furthermore, we extend our model by introducing new parameters and functions to influence the quantity of the topical features presented in the generated text. This feature would allow us to easily control the topical properties of the generated text. Our experimental results demonstrate that our model outperforms the state-of-the-art results on coherency, diversity, and fluency while being faster in decoding.
Designing Disaggregated Evaluations of AI Systems: Choices, Considerations, and Tradeoffs
Barocas, Solon, Guo, Anhong, Kamar, Ece, Krones, Jacquelyn, Morris, Meredith Ringel, Vaughan, Jennifer Wortman, Wadsworth, Duncan, Wallach, Hanna
Several pieces of work have uncovered performance disparities by conducting "disaggregated evaluations" of AI systems. We build on these efforts by focusing on the choices that must be made when designing a disaggregated evaluation, as well as some of the key considerations that underlie these design choices and the tradeoffs between these considerations. We argue that a deeper understanding of the choices, considerations, and tradeoffs involved in designing disaggregated evaluations will better enable researchers, practitioners, and the public to understand the ways in which AI systems may be underperforming for particular groups of people.
The 10 most innovative companies in artificial intelligence
In 2020, renowned semiconductor maker Arm launched two new chips that are designed to bring AI beyond smartphones and tablets. The chips, called the Arm Cortex-M55 and the Ethos-U55, can power machine learning on billions more devices and sensors within the internet of things. This is part of Arm's push into what's known as "TinyML," where computationally intensive software can run entirely on small, super low-powered devices without needing access to the internet or the cloud. This kind of AI is already showing up in small wearables like the Oura ring and in smart asthma inhalers, but it has other applications in farming, biometrics, and the smart home. Now, pending regulatory approval, these capabilities will be part of the Nvidia family: In September 2020, the U.S. chipmaker agreed to acquire Arm for $40 billion.
AI uncovers Eli Lilly's rheumatoid arthritis drug Olumiant as potential Alzheimer's treatment
Could janus kinase (JAK) inhibitors like Eli Lilly's rheumatoid arthritis drug Olumiant be repurposed to treat Alzheimer's disease? Researchers at Harvard University and Massachusetts General Hospital have set out to find the answer to that question with a new clinical trial that was born from artificial intelligence. The researchers used a type of AI called machine learning to identify existing drugs that might be able to prevent neuronal death in Alzheimer's. The screen pulled up a list of 15 FDA-approved drugs as candidates for repurposing in Alzheimer's, and five of them were JAK inhibitors, they reported in the journal Nature Communications. JAK proteins fuel inflammation and have long been suspected to play a role in Alzheimer's.
NASA Perseverance rover checks its robotic arm on Mars in new photos
The pictures were shared through the'RAW images' feed on the NASA Perseverance website, which showcases every grab from every camera on the SUV-sized rover NASA's mission will search for signs of ancient life on on the Red Planet. Named Perseverance, the main car-sized rover will explore an ancient river delta within the Jezero Crater. This was once filled with a 1,600ft deep lake and may have been flooded multiple times when Mars was warm. It is believed the region hosted microbial life some 3.5 to 3.9 billion years ago and the rover will examine soil samples to hunt for proof. Perseverance landed inside the crater on February 18 and will also collect samples of the soil to return to Earth.