Materials
Partitioned Active Learning for Heterogeneous Systems
Lee, Cheolhei, Wang, Kaiwen, Wu, Jianguo, Cai, Wenjun, Yue, Xiaowei
Active learning is a subfield of machine learning that focuses on improving the data collection efficiency of expensive-to-evaluate systems. Especially, active learning integrated surrogate modeling has shown remarkable performance in computationally demanding engineering systems. However, the existence of heterogeneity in underlying systems may adversely affect the performance of active learning. In order to improve the learning efficiency under this regime, we propose the partitioned active learning that seeks the most informative design points for partitioned Gaussian process modeling of heterogeneous systems. The proposed active learning consists of two systematic subsequent steps: the global searching scheme accelerates the exploration of active learning by investigating the most uncertain design space, and the local searching exploits the circumscribed information induced by the local GP. We also propose Cholesky update driven numerical remedies for our active learning to address the computational complexity challenge. The proposed method is applied to numerical simulations and two real-world case studies about (i) the cost-efficient automatic fuselage shape control in aerospace manufacturing; and (ii) the optimal design of tribocorrosion-resistant alloys in materials science. The results show that our approach outperforms benchmark methods with respect to prediction accuracy and computational efficiency.
cgSpan: Pattern Mining in Conceptual Graphs
Faci, Adam, Lesot, Marie-Jeanne, Laudy, Claire
Conceptual Graphs (CGs) are a graph-based knowledge representation formalism. In this paper we propose cgSpan a CG frequent pattern mining algorithm. It extends the DMGM-GSM algorithm that takes taxonomy-based labeled graphs as input; it includes three more kinds of knowledge of the CG formalism: (a) the fixed arity of relation nodes, handling graphs of neighborhoods centered on relations rather than graphs of nodes, (b) the signatures, avoiding patterns with concept types more general than the maximal types specified in signatures and (c) the inference rules, applying them during the pattern mining process. The experimental study highlights that cgSpan is a functional CG Frequent Pattern Mining algorithm and that including CGs specificities results in a faster algorithm with more expressive results and less redundancy with vocabulary.
Hit and Lead Discovery with Explorative RL and Fragment-based Molecule Generation
Yang, Soojung, Hwang, Doyeong, Lee, Seul, Ryu, Seongok, Hwang, Sung Ju
Recently, utilizing reinforcement learning (RL) to generate molecules with desired properties has been highlighted as a promising strategy for drug design. A molecular docking program - a physical simulation that estimates protein-small molecule binding affinity - can be an ideal reward scoring function for RL, as it is a straightforward proxy of the therapeutic potential. Still, two imminent challenges exist for this task. First, the models often fail to generate chemically realistic and pharmacochemically acceptable molecules. Second, the docking score optimization is a difficult exploration problem that involves many local optima and less smooth surfaces with respect to molecular structure. To tackle these challenges, we propose a novel RL framework that generates pharmacochemically acceptable molecules with large docking scores. Our method - Fragment-based generative RL with Explorative Experience replay for Drug design (FREED) - constrains the generated molecules to a realistic and qualified chemical space and effectively explores the space to find drugs by coupling our fragment-based generation method and a novel error-prioritized experience replay (PER). We also show that our model performs well on both de novo and scaffold-based schemes. Our model produces molecules of higher quality compared to existing methods while achieving state-of-the-art performance on two of three targets in terms of the docking scores of the generated molecules. We further show with ablation studies that our method, predictive error-PER (FREED(PE)), significantly improves the model performance.
Apple selects Chinese giant for critical iPhone role - California News Times
This article is an on-site version of the #techAsia newsletter.sign up here Send newsletter directly to your inbox every Wednesday Hello, Kenji from Tokyo this week is currently undergoing home quarantine for Covid-19. For our big story, there is another scoop about Apple from Nikkei Asia. China's state-owned enterprise has become a supplier of the latest flagship iPhone displays. This shows how advanced China's technology, including artificial intelligence, has advanced, as warned by a former Pentagon chief software officer (Mercedes Top 10). Meanwhile, China is building and diversifying its sources of strategic mineral resources, including lithium, a key component of the world's leading electric vehicle industry (our views, smart data and spotlights).
How to Improve Deep Learning Forecasts for Time Series
Clustering time series data before fitting can improve accuracy by 33% -- src. In 2021, researchers at UCLA developed a method that can improve model fit on many different time series'. By aggregating similarly structured data and fitting a model to each group, our models can specialize. While fairly straightforward to implement, as with any other complex deep learning method, we are often computationally limited by large data sets. However, all of the methods listed have support in both R and python, so development on smaller datasets should be pretty "simple."
How Moveworks' AI platform broke through the multilingual NLP barrier
Chatbots have a checkered past of often not delivering the performance their providers have promised. This is especially true in the IT service management (ITSM) and multilingual NLP spaces, where service desks found support teams deluged with complaints -- yes, about the support chatbots. Just getting English language nuance right and how enterprises communicate often require chatbots to be custom programmed with constraint and logic workflows supported with natural language processing (NLP) and machine learning. If that sounds like a science project, it is, and IT users are the test subjects. Because of their complexity, chatbots were contributing to already overflowing trouble-ticket queues.
Applications and Techniques for Fast Machine Learning in Science
Deiana, Allison McCarn, Tran, Nhan, Agar, Joshua, Blott, Michaela, Di Guglielmo, Giuseppe, Duarte, Javier, Harris, Philip, Hauck, Scott, Liu, Mia, Neubauer, Mark S., Ngadiuba, Jennifer, Ogrenci-Memik, Seda, Pierini, Maurizio, Aarrestad, Thea, Bahr, Steffen, Becker, Jurgen, Berthold, Anne-Sophie, Bonventre, Richard J., Bravo, Tomas E. Muller, Diefenthaler, Markus, Dong, Zhen, Fritzsche, Nick, Gholami, Amir, Govorkova, Ekaterina, Hazelwood, Kyle J, Herwig, Christian, Khan, Babar, Kim, Sehoon, Klijnsma, Thomas, Liu, Yaling, Lo, Kin Ho, Nguyen, Tri, Pezzullo, Gianantonio, Rasoulinezhad, Seyedramin, Rivera, Ryan A., Scholberg, Kate, Selig, Justin, Sen, Sougata, Strukov, Dmitri, Tang, William, Thais, Savannah, Unger, Kai Lukas, Vilalta, Ricardo, Krosigk, Belinavon, Warburton, Thomas K., Flechas, Maria Acosta, Aportela, Anthony, Calvet, Thomas, Cristella, Leonardo, Diaz, Daniel, Doglioni, Caterina, Galati, Maria Domenica, Khoda, Elham E, Fahim, Farah, Giri, Davide, Hawks, Benjamin, Hoang, Duc, Holzman, Burt, Hsu, Shih-Chieh, Jindariani, Sergo, Johnson, Iris, Kansal, Raghav, Kastner, Ryan, Katsavounidis, Erik, Krupa, Jeffrey, Li, Pan, Madireddy, Sandeep, Marx, Ethan, McCormack, Patrick, Meza, Andres, Mitrevski, Jovan, Mohammed, Mohammed Attia, Mokhtar, Farouk, Moreno, Eric, Nagu, Srishti, Narayan, Rohin, Palladino, Noah, Que, Zhiqiang, Park, Sang Eon, Ramamoorthy, Subramanian, Rankin, Dylan, Rothman, Simon, Sharma, Ashish, Summers, Sioni, Vischia, Pietro, Vlimant, Jean-Roch, Weng, Olivia
In this community review report, we discuss applications and techniques for fast machine learning (ML) in science -- the concept of integrating power ML methods into the real-time experimental data processing loop to accelerate scientific discovery. The material for the report builds on two workshops held by the Fast ML for Science community and covers three main areas: applications for fast ML across a number of scientific domains; techniques for training and implementing performant and resource-efficient ML algorithms; and computing architectures, platforms, and technologies for deploying these algorithms. We also present overlapping challenges across the multiple scientific domains where common solutions can be found. This community report is intended to give plenty of examples and inspiration for scientific discovery through integrated and accelerated ML solutions. This is followed by a high-level overview and organization of technical advances, including an abundance of pointers to source material, which can enable these breakthroughs.
Can It Really Do That? -- Introducing the Edge X AI Camera
The MXC Foundation has made a remarkable entry into the nascent multi-billion dollar AI smart device market. With the exponential growth of its network across the globe, the Foundation is thrilled to introduce more aspects to its network usage, allowing its mining community to utilize the data republic and see the network in action. The proprietary MXProtocol, together with scalable and secure aspects of device provisioning that connect with sensor technology, has proven successful and brings us a step closer to realizing truly smart cities. One such use case, which the MXC Foundation recently tested in a controlled environment, was the Edge X AI Camera. Read on to find out more about all the great functionalities packed into one small device.
MIT accelerates the discovery of new 3D printing materials with open-source AI platform
A partnership between the Massachusetts Institute of Technology and the chemical giant BASF has managed to successfully create an AI-driven process to speed up the discovery of custom 3D printing materials. Chemists usually develop a few iterations of a material candidate over a couple of days and test them in the lab. The new machine-learning algorithm can churn out hundreds of those iterations with the desired characteristics in the same timeframe. This would save time and raw material costs, as well as lessen the environmental impact of the discarded chemicals. Not only that, but the algorithm may also come up with ideas that the material's engineer could have overlooked for various reasons.
MIT Uses AI To Accelerate the Discovery of New Materials for 3D Printing
Researchers at MIT and BASF have developed a data-driven system that accelerates the process of discovering new 3D printing materials that have multiple mechanical properties. A new machine-learning system costs less, generates less waste, and can be more innovative than manual discovery methods. The growing popularity of 3D printing for manufacturing all sorts of items, from customized medical devices to affordable homes, has created more demand for new 3D printing materials designed for very specific uses. To cut down on the time it takes to discover these new materials, researchers at MIT have developed a data-driven process that uses machine learning to optimize new 3D printing materials with multiple characteristics, like toughness and compression strength. By streamlining materials development, the system lowers costs and lessens the environmental impact by reducing the amount of chemical waste.