Regression
Machine Learning to Estimate Gross Loss of Jewelry for Wax Patterns
Jain, Mihir, Jain, Kashish, Mane, Sandip
In mass manufacturing of jewellery, the gross loss is estimated before manufacturing to calculate the wax weight of the pattern that would be investment casted to make multiple identical pieces of jewellery. Machine learning is a technology that is a part of AI which helps create a model with decision-making capabilities based on a large set of user-defined data. In this paper, the authors found a way to use Machine Learning in the jewellery industry to estimate this crucial Gross Loss. Choosing a small data set of manufactured rings and via regression analysis, it was found out that there is a potential of reducing the error in estimation from +-2-3 to +-0.5 using ML Algorithms from historic data and attributes collected from the CAD file during the design phase itself. To evaluate the approach's viability, additional study must be undertaken with a larger data set.
How Bayesian additive regression trees(BART) are used part2(Machine Learning)
Abstract: Methods utilizing instrumental variables have been a fundamental statistical approach to estimation in the presence of unmeasured confounding, usually occurring in non-randomized observational data common to fields such as economics and public health. However, such methods usually make constricting linearity and additivity assumptions that are inapplicable to the complex modeling challenges of today. The growing body of observational data being collected will necessitate flexible regression modeling while also being able to control for confounding using instrumental variables. Therefore, this article presents a nonlinear instrumental variable regression model based on Bayesian regression tree ensembles to estimate such relationships, including interactions, in the presence of confounding. One exciting application of this method is to use genetic variants as instruments, known as Mendelian randomization.
How Bayesian additive regression trees(BART) are used part3(Machine Learning)
Abstract: Using ensemble methods for regression has been a large success in obtaining high-accuracy prediction. Examples are Bagging, Random forest, Boosting, BART (Bayesian additive regression tree), and their variants. In this paper, we propose a new perspective named variable grouping to enhance the predictive performance. The main idea is to seek for potential grouping of variables in such way that there is no nonlinear interaction term between variables of different groups. Given a sum-of-learner model, each learner will only be responsible for one group of variables, which would be more efficient in modeling nonlinear interactions.
Isotonic Recalibration under a Low Signal-to-Noise Ratio
Wรผthrich, Mario V., Ziegel, Johanna
There are two seemingly unrelated problems in insurance pricing that we are going to tackle in this paper. First, an insurance pricing system should not have any systematic cross-financing between different price cohorts. Systematic cross-financing implicitly means that some parts of the portfolio are under-priced, and this is compensated by other parts of the portfolio that are over-priced. We can prevent systematic cross-financing between price cohorts by ensuring that the pricing system is auto-calibrated. We propose to apply isotonic recalibration which turns any regression function into an auto-calibrated pricing system.
Large-Scale Cell-Level Quality of Service Estimation on 5G Networks Using Machine Learning Techniques
ฤฐลyapar, M. Tuฤberk, Uyan, Ufuk, รztรผrk, Mahiye Uluyaฤmur
This study presents a general machine learning framework to estimate the traffic-measurement-level experience rate at given throughput values in the form of a Key Performance Indicator for the cells on base stations across various cities, using busy-hour counter data, and several technical parameters together with the network topology. Relying on feature engineering techniques, scores of additional predictors are proposed to enhance the effects of raw correlated counter values over the corresponding targets, and to represent the underlying interactions among groups of cells within nearby spatial locations effectively. An end-to-end regression modeling is applied on the transformed data, with results presented on unseen cities of varying sizes.
Plant species richness prediction from DESIS hyperspectral data: A comparison study on feature extraction procedures and regression models
Guo, Yiqing, Mokany, Karel, Ong, Cindy, Moghadam, Peyman, Ferrier, Simon, Levick, Shaun R.
The diversity of terrestrial vascular plants plays a key role in maintaining the stability and productivity of ecosystems. Monitoring species compositional diversity across large spatial scales is challenging and time consuming. The advanced spectral and spatial specification of the recently launched DESIS (the DLR Earth Sensing Imaging Spectrometer) instrument provides a unique opportunity to test the potential for monitoring plant species diversity with spaceborne hyperspectral data. This study provides a quantitative assessment on the ability of DESIS hyperspectral data for predicting plant species richness in two different habitat types in southeast Australia. Spectral features were first extracted from the DESIS spectra, then regressed against on-ground estimates of plant species richness, with a two-fold cross validation scheme to assess the predictive performance. We tested and compared the effectiveness of Principal Component Analysis (PCA), Canonical Correlation Analysis (CCA), and Partial Least Squares analysis (PLS) for feature extraction, and Kernel Ridge Regression (KRR), Gaussian Process Regression (GPR), Random Forest Regression (RFR) for species richness prediction. The best prediction results were r=0.76 and RMSE=5.89 for the Southern Tablelands region, and r=0.68 and RMSE=5.95 for the Snowy Mountains region. Relative importance analysis for the DESIS spectral bands showed that the red-edge, red, and blue spectral regions were more important for predicting plant species richness than the green bands and the near-infrared bands beyond red-edge. We also found that the DESIS hyperspectral data performed better than Sentinel-2 multispectral data in the prediction of plant species richness. Our results provide a quantitative reference for future studies exploring the potential of spaceborne hyperspectral data for plant biodiversity mapping.
Sequentially Controlled Text Generation
Spangher, Alexander, Hua, Xinyu, Ming, Yao, Peng, Nanyun
While GPT-2 generates sentences that are remarkably human-like, longer documents can ramble and do not follow human-like writing structure. We study the problem of imposing structure on long-range text. We propose a novel controlled text generation task, sequentially controlled text generation, and identify a dataset, NewsDiscourse as a starting point for this task. We develop a sequential controlled text generation pipeline with generation and editing. We test different degrees of structural awareness and show that, in general, more structural awareness results in higher control-accuracy, grammaticality, coherency and topicality, approaching human-level writing performance.
One-vs-All Logistic Regression for Image Recognition in Python
This article represents the continuation of a series of dedicated articles that began some time ago. This series proposes the reader to understand the basic concepts leading to Machine Learning for biomedical data, like the difference between Linear and logistic regression, the Cost Function, Regularized Logistic Regression, and Gradient (see the Reference section). Each implementation is intended from scratch, and we will not use optimized machine learning packages like Scikit-learn, PyTorch, or TensorFlow. The only requirement is an updated version of Python 3, some fundamental libraries, and the desire to read this post to the end! Regressions (linear, logistic, for single and multiple variables) are statistical models helpful in finding correlations between observed dataset variables and answering whether those correlations are statistically significant.
$l_{1-2}$ GLasso: $L_{1-2}$ Regularized Multi-task Graphical Lasso for Joint Estimation of eQTL Mapping and Gene Network
Developments in sequencing technology allow us to obtain more and more genomic data since the publication of the first human genome sequence. Computational techniques can help us to mine meaningful information from raw data and understand how gene expression is regulated in cells. In general, these problems include identifying cancer gene co-expression (co-expression: simultaneous expression of two or more genes) modules, determining SNP-gene relationships through eQTL (expression quantitative trait locus) mapping and determining gene-gene relationships by estimating gene network structure, etc (Rockman and Kruglyak, 2006; Gardner and Faith, 2005). Given a dataset containing single nucleotide polymorphisms (SNPs) and mRNA expression, the problem is to understand the SNP-gene and gene-gene relationships.
Machine-Learning Prediction of the Computed Band Gaps of Double Perovskite Materials
Zhang, Junfei, Li, Yueqi, Zhou, Xinbo
Prediction of the electronic structure of functional materials is essential for the engineering of new devices. Conventional electronic structure prediction methods based on density functional theory (DFT) suffer from not only high computational cost, but also limited accuracy arising from the approximations of the exchange-correlation functional. Surrogate methods based on machine learning have garnered much attention as a viable alternative to bypass these limitations, especially in the prediction of solid-state band gaps, which motivated this research study. Herein, we construct a random forest regression model for band gaps of double perovskite materials, using a dataset of 1306 band gaps computed with the GLLBSC (Gritsenko, van Leeuwen, van Lenthe, and Baerends solid correlation) functional. Among the 20 physical features employed, we find that the bulk modulus, superconductivity temperature, and cation electronegativity exhibit the highest importance scores, consistent with the physics of the underlying electronic structure. Using the top 10 features, a model accuracy of 85.6% with a root mean square error of 0.64 eV is obtained, comparable to previous studies. Our results are significant in the sense that they attest to the potential of machine learning regressions for the rapid screening of promising candidate functional materials.