Materials
Predicting Quality of Wine with Data Science and Machine Learning - TechnologyHQ
Wine is the most widely consumed beverage globally, and its value is important to society. The quality of wine is significant to its consumers and producers in the current competitive market. Wine quality was determined historically by the testing done at the end. To achieve that level, one must spend a lot of money, time and follow the various procedures from the beginning to get good quality wine. Traditionally, this proved to be very expensive.
Searching for Structure in Unfalsifiable Claims
Christensen, Peter Ebert, Warburg, Frederik, Jia, Menglin, Belongie, Serge
Social media platforms give rise to an abundance of posts and comments on every topic imaginable. Many of these posts express opinions on various aspects of society, but their unfalsifiable nature makes them ill-suited to fact-checking pipelines. In this work, we aim to distill such posts into a small set of narratives that capture the essential claims related to a given topic. Understanding and visualizing these narratives can facilitate more informed debates on social media. As a first step towards systematically identifying the underlying narratives on social media, we introduce PAPYER, a fine-grained dataset of online comments related to hygiene in public restrooms, which contains a multitude of unfalsifiable claims. We present a human-in-the-loop pipeline that uses a combination of machine and human kernels to discover the prevailing narratives and show that this pipeline outperforms recent large transformer models and state-of-the-art unsupervised topic models.
Trustworthy modelling of atmospheric formaldehyde powered by deep learning
Biswas, Mriganka Sekhar, Singh, Manmeet
Formaldehyde (HCHO) is one one of the most important trace gas in the atmosphere, as it is a pollutant causing respiratory and other diseases. It is also a precursor of tropospheric ozone which damages crops and deteriorates human health. Study of HCHO chemistry and long-term monitoring using satellite data is important from the perspective of human health, food security and air pollution. Dynamic atmospheric chemistry models struggle to simulate atmospheric formaldehyde and often overestimate by up to two times relative to satellite observations and reanalysis. Spatial distribution of modelled HCHO also fail to match satellite observations. Here, we present deep learning approach using a simple super-resolution based convolutional neural network towards simulating fast and reliable atmospheric HCHO. Our approach is an indirect method of HCHO estimation without the need to chemical equations. We find that deep learning outperforms dynamical model simulations which involves complicated atmospheric chemistry representation. Causality establishing the nonlinear relationships of different variables to target formaldehyde is established in our approach by using a variety of precursors from meteorology and chemical reanalysis to target OMI AURA satellite based HCHO predictions. We choose South Asia for testing our implementation as it doesnt have in situ measurements of formaldehyde and there is a need for improved quality data over the region. Moreover, there are spatial and temporal data gaps in the satellite product which can be removed by trustworthy modelling of atmospheric formaldehyde. This study is a novel attempt using computer vision for trustworthy modelling of formaldehyde from remote sensing can lead to cascading societal benefits.
Towards Automated Process Planning and Mining
Fettke, Peter, Rombach, Alexander
AI Planning, Machine Learning and Process Mining have so far developed into separate research fields. At the same time, many interesting concepts and insights have been gained at the intersection of these areas in recent years. For example, the behavior of future processes is now comprehensively predicted with the aid of Machine Learning. For the practical application of these findings, however, it is also necessary not only to know the expected course, but also to give recommendations and hints for the achievement of goals, i.e. to carry out comprehensive process planning. At the same time, an adequate integration of the aforementioned research fields is still lacking. In this article, we present a research project in which researchers from the AI and BPM field work jointly together. Therefore, we discuss the overall research problem, the relevant fields of research and our overall research framework to automatically derive process models from executional process data, derive subsequent planning problems and conduct automated planning in order to adaptively plan and execute business processes using real-time forecasts.
Knowledge-Injected Federated Learning
Fan, Zhenan, Zhou, Zirui, Pei, Jian, Friedlander, Michael P., Hu, Jiajie, Li, Chengliang, Zhang, Yong
With the development of artificial intelligence, people recognize that many powerful machine learning models are driven by large decentralized datasets of various data types. However, in many industryscale applications, training data is obtained and maintained by different data owners instead of centralized at the data center, and sharing data is often forbidden due to privacy requirements. Federated learning (FL) is an emerging machine learning framework in which multiple data owners (also referred to as clients) participate in collaboratively training a model without sharing their local data with each other [18, 33]. Another challenge with artificial intelligence is integrating domain knowledge into purely datadriven models, i.e., parameters of the model are learned through training data without any human engineering [8, 11]. For example, human know-how and craftsmanship, which may not be learnable from the training data, can be formulated as prediction models, and combing them with a purely data-driven model may boost its performance and reduce the risk of overfitting [10].
Self-Organizing Map Neural Network Algorithm for the Determination of Fracture Location in Solid-State Process joined Dissimilar Alloys
Mishra, Akshansh, Dasgupta, Anish
The philosophical movement known as computational mind theory or computationalism, which promotes the idea that neural computation accounts cognition, has ties to neural computation [1-4]. Nowadays, these types of algorithms are used in manufacturing and materials sectors for the determination of mechanical and microstructure properties of fabricated alloys or specimens [5-6]. An artificial neural network (ANN) was used by Shiau et al. [7] to model Taiwan's industrial energy demand in relation to subsector industrial output and climate change. It was the first investigation to measure the relationship between industrial energy use, manufacturing output, and climate change using the ANN technique. A multilayer perceptron (MLP) with a feedforward backpropagation neural network was used as the ANN model in this investigation. In order to improve the implementation of natural fibers in green bio-composites, Jarrah et al. [8] used doubly interconnected artificial neural networks to make unique classifications and prediction of the inherent mechanical properties of natural fibers.
Towards out of distribution generalization for problems in mechanics
Yuan, Lingxiao, Park, Harold S., Lejeune, Emma
There has been a massive increase in research interest towards applying data driven methods to problems in mechanics. While traditional machine learning (ML) methods have enabled many breakthroughs, they rely on the assumption that the training (observed) data and testing (unseen) data are independent and identically distributed (i.i.d). Thus, traditional ML approaches often break down when applied to real world mechanics problems with unknown test environments and data distribution shifts. In contrast, out-of-distribution (OOD) generalization assumes that the test data may shift (i.e., violate the i.i.d. assumption). To date, multiple methods have been proposed to improve the OOD generalization of ML methods. However, because of the lack of benchmark datasets for OOD regression problems, the efficiency of these OOD methods on regression problems, which dominate the mechanics field, remains unknown. To address this, we investigate the performance of OOD generalization methods for regression problems in mechanics. Specifically, we identify three OOD problems: covariate shift, mechanism shift, and sampling bias. For each problem, we create two benchmark examples that extend the Mechanical MNIST dataset collection, and we investigate the performance of popular OOD generalization methods on these mechanics-specific regression problems. Our numerical experiments show that in most cases, while the OOD generalization algorithms perform better compared to traditional ML methods on these OOD problems, there is a compelling need to develop more robust OOD generalization methods that are effective across multiple OOD scenarios. Overall, we expect that this study, as well as the associated open access benchmark datasets, will enable further development of OOD generalization methods for mechanics specific regression problems.
Incoporating Weighted Board Learning System for Accurate Occupational Pneumoconiosis Staging
Yang, Kaiguang, Wang, Yeping, Luo, Qianhao, Liu, Xin, Li, Weiling
Occupational pneumoconiosis (OP) staging is a vital task concerning the lung healthy of a subject. The staging result of a patient is depended on the staging standard and his chest X-ray. It is essentially an image classification task. However, the distribution of OP data is commonly imbalanced, which largely reduces the effect of classification models which are proposed under the assumption that data follow a balanced distribution and causes inaccurate staging results. To achieve accurate OP staging, we proposed an OP staging model who is able to handle imbalance data in this work. The proposed model adopts gray level co-occurrence matrix (GLCM) to extract texture feature of chest X-ray and implements classification with a weighted broad learning system (WBLS). Empirical studies on six data cases provided by a hospital indicate that proposed model can perform better OP staging than state-of-the-art classifiers with imbalanced data.
How AI could fuel global warming
Data and cloud are not virtual technology. They need costly infrastructure and electricity. Researchers forecast that in the future their emissions could be much more than expected. Who does not make us sleep the night? A few weeks ago the UK break the temperature record, for the first time the temperature rose over 40 C. The summer nights are warm and humid and it is hard to sleep on similar days.
Robotic Inspection and Characterization of Subsurface Defects on Concrete Structures Using Impact Sounding
Hoxha, Ejup, Feng, Jinglun, Sanakov, Diar, Gjinofci, Ardian, Xiao, Jizhong
Impact-sounding (IS) and impact-echo (IE) are well-developed non-destructive evaluation (NDE) methods that are widely used for inspections of concrete structures to ensure the safety and sustainability. However, it is a tedious work to collect IS and IE data along grid lines covering a large target area for characterization of subsurface defects. On the other hand, data processing is very complicated that requires domain experts to interpret the results. To address the above problems, we present a novel robotic inspection system named as Impact-Rover to automate the data collection process and introduce data analytics software to visualize the inspection result allowing regular non-professional people to understand. The system consists of three modules: 1) a robotic platform with vertical mobility to collect IS and IE data in hard-to-reach locations, 2) vision-based positioning module that fuses the RGB-D camera, IMU and wheel encoder to estimate the 6-DOF pose of the robot, 3) a data analytics software module for processing the IS data to generate defect maps. The Impact-Rover hosts both IE and IS devices on a sliding mechanism and can perform move-stop-sample operations to collect multiple IS and IE data at adjustable spacing. The robot takes samples much faster than the manual data collection method because it automatically takes the multiple measurements along a straight line and records the locations. This paper focuses on reporting experimental results on IS. We calculate features and use unsupervised learning methods for analyzing the data. By combining the pose generated by our vision-based localization module and the position of the head of the sliding mechanism we can generate maps of possible defects. The results on concrete slabs demonstrate that our impact-sounding system can effectively reveal shallow defects.