Goto

Collaborating Authors

 Regression


AI/ML Algorithms and Applications in VLSI Design and Technology

arXiv.org Artificial Intelligence

An evident challenge ahead for the integrated circuit (IC) industry in the nanometer regime is the investigation and development of methods that can reduce the design complexity ensuing from growing process variations and curtail the turnaround time of chip manufacturing. Conventional methodologies employed for such tasks are largely manual; thus, time-consuming and resource-intensive. In contrast, the unique learning strategies of artificial intelligence (AI) provide numerous exciting automated approaches for handling complex and data-intensive tasks in very-large-scale integration (VLSI) design and testing. Employing AI and machine learning (ML) algorithms in VLSI design and manufacturing reduces the time and effort for understanding and processing the data within and across different abstraction levels via automated learning algorithms. It, in turn, improves the IC yield and reduces the manufacturing turnaround time. This paper thoroughly reviews the AI/ML automated approaches introduced in the past towards VLSI design and manufacturing. Moreover, we discuss the scope of AI/ML applications in the future at various abstraction levels to revolutionize the field of VLSI design, aiming for high-speed, highly intelligent, and efficient implementations.


Contrastive Learning Can Find An Optimal Basis For Approximately View-Invariant Functions

arXiv.org Artificial Intelligence

Contrastive learning is a powerful framework for learning self-supervised representations that generalize well to downstream supervised tasks. We show that multiple existing contrastive learning methods can be reinterpreted as learning kernel functions that approximate a fixed positive-pair kernel. We then prove that a simple representation obtained by combining this kernel with PCA provably minimizes the worst-case approximation error of linear predictors, under a straightforward assumption that positive pairs have similar labels. Our analysis is based on a decomposition of the target function in terms of the eigenfunctions of a positive-pair Markov chain, and a surprising equivalence between these eigenfunctions and the output of Kernel PCA. We give generalization bounds for downstream linear prediction using our Kernel PCA representation, and show empirically on a set of synthetic tasks that applying Kernel PCA to contrastive learning models can indeed approximately recover the Markov chain eigenfunctions, although the accuracy depends on the kernel parameterization as well as on the augmentation strength.


Joint Probability Trees

arXiv.org Artificial Intelligence

Joint probability distributions offer a wide range of highpotential applications in engineering, science, and technology (Chater et al., 2006; Griffiths et al., 2008; Knill & Pouget, 2004). Besides families of continuous distributions, introducing strong independence assumptions that must be probabilistic graphical models (PGMs), such as Bayesian known prior to learning and may turn out to be too great networks and Markov random fields (Koller & Friedman, simplifications of a model to be of practical use (Besag, 2009), are the de-facto standard in probabilistic knowledge 1975; Jain, 2012). As a simple example, consider a probability representation. They provide graph-based languages to space X, Y, C of two numeric variables, X and Y, model dependencies and independencies of variables, and and one symbolic variable C, dom(C) = {Red, Blue} as local joint or conditional distributions that quantify the statistical illustrated in Figures 1a and 1b. Let the symbolic values Red dependencies. However, the practical applicability of and Blue demaracate two clusters that are approximately PGMs suffers from the representational and computational normally distributed.


A model-free feature selection technique of feature screening and random forest based recursive feature elimination

arXiv.org Artificial Intelligence

Due to the development of data technology, feature selection plays an important role in both statistics and machine learning. High dimensional and ultra-high dimensional datasets are widely used in many fields, such as finance, image recognition, text classification, etc. Although more detailed information is provided with the increase of dimensions, the existence of a large number of redundant features weakens the generalization ability of models and increases the difficulty of data analysis (Jain et al., 2000). Thus, the efficiency of feature selection is crucial, as it focuses on choosing a small subset of informative features that contain the information of data and determine the source of the specific concerns derived from the study. In many data analyses, feature selection is a significant and frequently used dimensionality reduction technique and is often considered a key preprocessing step in data analysis for model benefits, such as interpretability, accuracy, lower computational costs, and less prone to overfitting (Sheikhpour et al., 2017; Khaire and Dhanalakshmi, 2022).


Provable Detection of Propagating Sampling Bias in Prediction Models

arXiv.org Artificial Intelligence

With an increased focus on incorporating fairness in machine learning models, it becomes imperative not only to assess and mitigate bias at each stage of the machine learning pipeline but also to understand the downstream impacts of bias across stages. Here we consider a general, but realistic, scenario in which a predictive model is learned from (potentially biased) training data, and model predictions are assessed post-hoc for fairness by some auditing method. We provide a theoretical analysis of how a specific form of data bias, differential sampling bias, propagates from the data stage to the prediction stage. Unlike prior work, we evaluate the downstream impacts of data biases quantitatively rather than qualitatively and prove theoretical guarantees for detection. Under reasonable assumptions, we quantify how the amount of bias in the model predictions varies as a function of the amount of differential sampling bias in the data, and at what point this bias becomes provably detectable by the auditor. Through experiments on two criminal justice datasets -- the well-known COMPAS dataset and historical data from NYPD's stop and frisk policy -- we demonstrate that the theoretical results hold in practice even when our assumptions are relaxed.


Masked Multi-Step Probabilistic Forecasting for Short-to-Mid-Term Electricity Demand

arXiv.org Artificial Intelligence

Predicting the demand for electricity with uncertainty helps in planning and operation of the grid to provide reliable supply of power to the consumers. Machine learning (ML)-based demand forecasting approaches can be categorized into (1) sample-based approaches, where each forecast is made independently, and (2) time series regression approaches, where some historical load and other feature information is used. When making a short-to-mid-term electricity demand forecast, some future information is available, such as the weather forecast and calendar variables. However, in existing forecasting models this future information is not fully incorporated. To overcome this limitation of existing approaches, we propose Masked Multi-Step Multivariate Probabilistic Forecasting (MMMPF), a novel and general framework to train any neural network model capable of generating a sequence of outputs, that combines both the temporal information from the past and the known information about the future to make probabilistic predictions. Experiments are performed on a real-world dataset for short-to-mid-term electricity demand forecasting for multiple regions and compared with various ML methods. They show that the proposed MMMPF framework outperforms not only sample-based methods but also existing time-series forecasting models with the exact same base models. Models trainded with MMMPF can also generate desired quantiles to capture uncertainty and enable probabilistic planning for grid of the future.


Low-dimensional Data-based Surrogate Model of a Continuum-mechanical Musculoskeletal System Based on Non-intrusive Model Order Reduction

arXiv.org Artificial Intelligence

In recent decades, the main focus of computer modeling has been on supporting the design and development of engineering prototyes, but it is now ubiquitous in non-traditional areas such as medical rehabilitation. Conventional modeling approaches like the finite element~(FE) method are computationally costly when dealing with complex models, making them of limited use for purposes like real-time simulation or deployment on low-end hardware, if the model at hand cannot be simplified in a useful manner. Consequently, non-traditional approaches such as surrogate modeling using data-driven model order reduction are used to make complex high-fidelity models more widely available anyway. They often involve a dimensionality reduction step, in which the high-dimensional system state is transformed onto a low-dimensional subspace or manifold, and a regression approach to capture the reduced system behavior. While most publications focus on one dimensionality reduction, such as principal component analysis~(PCA) (linear) or autoencoder (nonlinear), we consider and compare PCA, kernel PCA, autoencoders, as well as variational autoencoders for the approximation of a structural dynamical system. In detail, we demonstrate the benefits of the surrogate modeling approach on a complex FE model of a human upper-arm. We consider both the models deformation and the internal stress as the two main quantities of interest in a FE context. By doing so we are able to create a computationally low cost surrogate model which captures the system behavior with high approximation quality and fast evaluations.


Generalization Ability of Wide Neural Networks on $\mathbb{R}$

arXiv.org Artificial Intelligence

Deep neural networks have been successfully applied in various fields such as image analysis, natural language processing, protein structure prediction, etc.[40, 22, 35]. Since the number of parameters appeared in deep neural networks is often ten times or hundred times larger than the sample size of data, the successes of neural network methods have challenged the traditional bias variances trade-off principle, one of the primary doctrines in the classical statistical learning theories [61]. For example, many influential experiments [9, 67, 8, 48, 7] suggested that if one trains a neural network till it overfits the data, the resulting network can still generalize well. This observation, often referred to as the "benign overfitting phenomenon" [4, 53, 26, 45], actually reshaped the landscape of the studies in neural networks. For example, some researchers built giant neural networks in practice which can easily achieve nearly zero training error and possess the state-of-the-art performances [31, 50, 21]. Inspired by these experiments and observations, researchers proposed various new theories to explain why overfitted neural networks do generalize well on certain data [9, 43, 26, 47]. Several groups of statisticians tried to explain the generalization ability of neural networks from statistical decision theory with various carefully designed nonparametric regression frameworks. For example, assuming that the regression function belongs to a carefully designed sub-class of the Hรถlder continuous functions, [5] proved that there exists a neural network with sigmoid activation function achieving the corresponding minimax rate; [54] further established similar results for ReLU neural networks based on the approximation theory from [66]; [59] then extended these results to regression functions in Besov space and its variants.


Deep Neural Networks for Nonparametric Interaction Models with Diverging Dimension

arXiv.org Machine Learning

Deep neural networks have achieved tremendous success due to their representation power and adaptation to low-dimensional structures. Their potential for estimating structured regression functions has been recently established in the literature. However, most of the studies require the input dimension to be fixed and consequently ignore the effect of dimension on the rate of convergence and hamper their applications to modern big data with high dimensionality. In this paper, we bridge this gap by analyzing a $k^{th}$ order nonparametric interaction model in both growing dimension scenarios ($d$ grows with $n$ but at a slower rate) and in high dimension ($d \gtrsim n$). In the latter case, sparsity assumptions and associated regularization are required in order to obtain optimal rates of convergence. A new challenge in diverging dimension setting is in calculation mean-square error, the covariance terms among estimated additive components are an order of magnitude larger than those of the variances and they can deteriorate statistical properties without proper care. We introduce a critical debiasing technique to amend the problem. We show that under certain standard assumptions, debiased deep neural networks achieve a minimax optimal rate both in terms of $(n, d)$. Our proof techniques rely crucially on a novel debiasing technique that makes the covariances of additive components negligible in the mean-square error calculation. In addition, we establish the matching lower bounds.


Hybrid Feature- and Similarity-Based Models for Joint Prediction and Interpretation

arXiv.org Artificial Intelligence

Electronic health records (EHRs) include simple features like patient age together with more complex data like care history that are informative but not easily represented as individual features. To better harness such data, we developed an interpretable hybrid feature- and similarity-based model for supervised learning that combines feature and kernel learning for prediction and for investigation of causal relationships. We fit our hybrid models by convex optimization with a sparsity-inducing penalty on the kernel. Depending on the desired model interpretation, the feature and kernel coefficients can be learned sequentially or simultaneously. The hybrid models showed comparable or better predictive performance than solely feature- or similarity-based approaches in a simulation study and in a case study to predict two-year risk of loneliness or social isolation with EHR data from a complex primary health care population. Using the case study we also present new kernels for high-dimensional indicator-coded EHR data that are based on deviations from population-level expectations, and we identify considerations for causal interpretations.