Goto

Collaborating Authors

 Country


Scale-Equivariant Neural Networks with Decomposed Convolutional Filters

arXiv.org Machine Learning

Encoding the input scale information explicitly into the representation learned by a convolutional neural network (CNN) is beneficial for many vision tasks especially when dealing with multiscale input signals. We study, in this paper, a scale-equivariant CNN architecture with joint convolutions across the space and the scaling group, which is shown to be both sufficient and necessary to achieve scale-equivariant representations. To reduce the model complexity and computational burden, we decompose the convolutional filters under two pre-fixed separable bases and truncate the expansion to low-frequency components. A further benefit of the truncated filter expansion is the improved deformation robustness of the equivariant representation. Numerical experiments demonstrate that the proposed scale-equivariant neural network with decomposed convolutional filters (ScDCFNet) achieves significantly improved performance in multiscale image classification and better interpretability than regular CNNs at a reduced model size.


Exascale Deep Learning for Scientific Inverse Problems

arXiv.org Machine Learning

We introduce novel communication strategies in synchronous distributed Deep Learning consisting of decentralized gradient reduction orchestration and computational graph-aware grouping of gradient tensors. Networks (DNN) models and data sets (Dai et al., 2019), the need for efficient distributed machine learning strategies on massively parallel systems is more significant than On small to moderate-scale systems, with 10's - 100's of GPU/TPU accelerators, these scaling inefficiencies can be difficult to detect and systematically optimize due to system noise and load variability. The scaling inefficiencies of data-parallel implementations are most readily apparent on large-scale systems such as supercomputers with 1,000's-10,000's of accelerators. Extending data-parallelism to the massive scale of super-computing systems is also motivated by the latter's traditional workload consisting of scientific numerical simulations (Kent & Kotliar, 2018). NVLink interconnect, supporting a (peak) bidirectional bandwidth of 100 GB/s, where each 3 V100 GPUs are grouped in a ring topology with all-to-all connections to a POWER9 CPU.


Supervised Vector Quantized Variational Autoencoder for Learning Interpretable Global Representations

arXiv.org Machine Learning

Learning interpretable representations of data remains a central challenge in deep learning. When training a deep generative model, the observed data are often associated with certain categorical labels, and, in parallel with learning to regenerate data and simulate new data, learning an interpretable representation of each class of data is also a process of acquiring knowledge. Here, we present a novel generative model, referred to as the Supervised Vector Quantized Variational AutoEncoder (S-VQ-VAE), which combines the power of supervised and unsupervised learning to obtain a unique, interpretable global representation for each class of data. Compared with conventional generative models, our model has three key advantages: first, it is an integrative model that can simultaneously learn a feature representation for individual data point and a global representation for each class of data; second, the learning of global representations with embedding codes is guided by supervised information, which clearly defines the interpretation of each code; and third, the global representations capture crucial characteristics of different classes, which reveal similarity and differences of statistical structures underlying different groups of data. We evaluated the utility of S-VQ-VAE on a machine learning benchmark dataset, the MNIST dataset, and on gene expression data from the Library of Integrated Network-Based Cellular Signatures (LINCS). We proved that S-VQ-VAE was able to learn the global genetic characteristics of samples perturbed by the same class of perturbagen (PCL), and further revealed the mechanism correlations between PCLs. Such knowledge is crucial for promoting new drug development for complex diseases like cancer.


A Neural Network Based Method to Solve Boundary Value Problems

arXiv.org Machine Learning

A Neural Network (NN) based numerical method is formulated and implemented for solving Boundary Value Problems (BVPs) and numerical results are presented to validate this method by solving Laplace equation with Dirichlet boundary condition and Poisson's equation with mixed boundary conditions. The principal advantage of NN based numerical method is the discrete data points where the field is computed, can be unstructured and do not suffer from issues of meshing like traditional numerical methods such as Finite Difference Time Domain or Finite Element Method. Numerical investigations are carried out for both uniform and non-uniform training grid distributions to understand the efficacy and limitations of this method and to provide qualitative understanding of various parameters involved.


Reservoir Topology in Deep Echo State Networks

arXiv.org Machine Learning

Deep Echo State Networks (DeepESNs) recently extended the applicability of Reservoir Computing (RC) methods towards the field of deep learning. In this paper we study the impact of constra ined reservoir topologies in the architectural design of deep reservo irs, through numerical experiments on several RC benchmarks. The major o utcome of our investigation is to show the remarkable effect, in term s of predictive performance gain, achieved by the synergy between a dee p reservoir construction and a structured organization of the recurren t units in each layer. Our results also indicate that a particularly advant ageous architectural setting is obtained in correspondence of DeepESNs whe re reservoir units are structured according to a permutation recurrent m atrix.


The column measure and Gradient-Free Gradient Boosting

arXiv.org Machine Learning

Sparse model selection by structural risk minimization leads to a set of a few predictors, ideally a subset of the true predictors. This selection clearly depends on the underlying loss function $\tilde L$. For linear regression with square loss, the particular (functional) Gradient Boosting variant $L_2-$Boosting excels for its computational efficiency even for very large predictor sets, while still providing suitable estimation consistency. For more general loss functions, functional gradients are not always easily accessible or, like in the case of continuous ranking, need not even exist. To close this gap, starting from column selection frequencies obtained from $L_2-$Boosting, we introduce a loss-dependent ''column measure'' $\nu^{(\tilde L)}$ which mathematically describes variable selection. The fact that certain variables relevant for a particular loss $\tilde L$ never get selected by $L_2-$Boosting is reflected by a respective singular part of $\nu^{(\tilde L)}$ w.r.t. $\nu^{(L_2)}$. With this concept at hand, it amounts to a suitable change of measure (accounting for singular parts) to make $L_2-$Boosting select variables according to a different loss $\tilde L$. As a consequence, this opens the bridge to applications of simulational techniques such as various resampling techniques, or rejection sampling, to achieve this change of measure in an algorithmic way.


IFR-Net: Iterative Feature Refinement Network for Compressed Sensing MRI

arXiv.org Machine Learning

To improve the compressive sensing MRI (CS - MRI) approaches in terms of fine structure loss under high acceleration factors, we have propose d an iterative feature refinement model (IFR - CS), equipped with fixed transforms, to restore the meaningful structure s and details. Nevertheless, the proposed IFR - CS still has some limitations, such as the selection of hyper - parameters, a lengthy reconstruction time, and the fixed sparsifying transform . To alleviate these issues, we unroll the iterative feature refinement procedure s in IFR - CS to a supervised model - driven network, dubbed IFR - Net. Equipped with training data pairs, both Additionally, inspired by the powerful representation capability of convolutional neural network (CNN), CNN - based inversion blocks are explored in the sparsity - promoting denoising module to generalize the sparsity - enforcing operator . Extensive experiments on both simulated and in v ivo MR datasets have shown that the proposed network possesses a strong capability to capture image details and preserve well the structural information with fast reconstruction speed. Index terms -- Compressed Sensing; Undersampled image reconstruction; IFR - CS; Deep learning; Model - driven network. Magnetic resonance imaging (MRI) is a non - invasive and widely used imaging technique that can provide both functional and anatomical information for clinic al diagnosis. However, the slow imaging speed may result in patient discomfort and motion artifacts. Therefore, increasing MR imaging speed is an important and worthwhile research goal. During the past decades, compressed sensing (CS) has become a popular and successful strategy for fast MR imaging reconstruction [1] - [6] . Zhang and Q. Yang are with the Department of Electronic Information Engineering, Nanchang Universi ty, Nanchang 330031, China. Liu did the work during her internship at Paul C. Lauterbur Research Center for Biomedical Imaging, Chinese Academy of Sciences, Shenzhen, China. S. Wang and D. Liang are with Paul C. Lauterbur Research Center for Biomedical Imaging and the Medical AI Research Center, Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences, Shenzhen 518055, China ( sophiasswang@hotmail.com, dong.liang@siat.ac.cn).


Entropy from Machine Learning

arXiv.org Machine Learning

Subsequently, one can use virtually any machine learning classification algorithm for computing entropy. This procedure can be used to compute entropy, and consequently the free energy directly from a set of Monte Carlo configurations at a given temperature. As a test of the proposed method, using an off-the-shelf machine learning classifier we reproduce the entropy and free energy of the 2D Ising model from Monte Carlo configurations at various temperatures throughout its phase diagram. Other potential applications include computing the entropy of spiking neurons or any other multidimensional binary signals. 1 Introduction The problem of estimating entropy of high dimensional binary configurations or signals is ubiquitous in many disciplines. In physics, we very often have at our disposal a set of configurations of some physical system generated by a Monte Carlo simulation at a given temperature T 0. This data is very much geared towards computing expectation values of various operators or their correlation functions, however obtaining the entropy or free energy of the system is far from trivial. Indeed, to the best of our knowledge, there is no known way to compute the entropy directly from these configurations even for a system of a quite moderate size (e.g. for a 20 20 lattice). The goal of the present paper is to propose machine email: romuald.janik@gmail.com


Understanding and Improving One-shot Neural Architecture Optimization

arXiv.org Machine Learning

The ability of accurately ranking candidate architectures is the key to the performance of neural architecture search~(NAS). One-shot NAS is proposed to cut the expense but shows inferior performance against conventional NAS and is not adequately stable. We find that the ranking correlation between architectures under one-shot training and the ones under stand-alone training is poor, which misleads the algorithm to discover better architectures. We conjecture that this is owing to the gaps between one-shot training and stand-alone complete training. In this work, we empirically investigate several main factors that lead to the gaps and so weak ranking correlation. We then propose NAO-V2 to alleviate such gaps where we: (1) Increase the average updates for individual architecture to a relatively adequate extent. (2) Encourage more updates for large and complex architectures than small and simple architectures to balance them by sampling architectures in proportion to their model sizes. (3) Make the one-shot training of the supernet independent at each iteration. Comprehensive experiments verify that our proposed method is effective and robust. It leads to a more stable search that all the top architectures perform well enough compared to baseline methods. The final discovered architecture shows significant improvements against baselines with a test error rate of 2.60% on CIFAR-10 and top-1 accuracy of 74.4% on ImageNet under the mobile setting. Code and model checkpoints are publicly available at https://github.com/renqianluo/NAO_pytorch.


WATTNet: Learning to Trade FX via Hierarchical Spatio-Temporal Representation of Highly Multivariate Time Series

arXiv.org Machine Learning

Finance is a particularly challenging application area for deep learning models due to low noise-to-signal ratio, non-stationarity, and partial observability. Non-deliverable-forwards (NDF), a derivatives contract used in foreign exchange (FX) trading, presents additional difficulty in the form of long-term planning required for an effective selection of start and end date of the contract. In this work, we focus on tackling the problem of NDF tenor selection by leveraging high-dimensional sequential data consisting of spot rates, technical indicators and expert tenor patterns. To this end, we construct a dataset from the Depository Trust & Clearing Corporation (DTCC) NDF data that includes a comprehensive list of NDF volumes and daily spot rates for 64 FX pairs. We introduce WaveATTentionNet (WATTNet), a novel temporal convolution (TCN) model for spatio-temporal modeling of highly multivariate time series, and validate it across NDF markets with varying degrees of dissimilarity between the training and test periods in terms of volatility and general market regimes. The proposed method achieves a significant positive return on investment (ROI) in all NDF markets under analysis, outperforming recurrent and classical baselines by a wide margin. Finally, we propose two orthogonal interpretability approaches to verify noise stability and detect the driving factors of the learned tenor selection strategy.