Deep Learning
Dynamic Knowledge embedding and tracing
Xu, Liangbei, Davenport, Mark A.
The goal of knowledge tracing is to track the state of a student's knowledge as it evolves over time. This plays a fundamental role in understanding the learning process and is a key task in the development of an intelligent tutoring system. In this paper we propose a novel approach to knowledge tracing that combines techniques from matrix factorization with recent progress in recurrent neural networks (RNNs) to effectively track the state of a student's knowledge. The proposed \emph{DynEmb} framework enables the tracking of student knowledge even without the concept/skill tag information that other knowledge tracing models require while simultaneously achieving superior performance. We provide experimental evaluations demonstrating that DynEmb achieves improved performance compared to baselines and illustrating the robustness and effectiveness of the proposed framework. We also evaluate our approach using several real-world datasets showing that the proposed model outperforms the previous state-of-the-art. These results suggest that combining embedding models with sequential models such as RNNs is a promising new direction for knowledge tracing.
Automatic Hip Fracture Identification and Functional Subclassification with Deep Learning
To investigate the feasibility of automatic identification and classification of hip fractures using deep learning, which may improve outcomes by reducing diagnostic errors and decreasing time to operation. Hip and pelvic radiographs from 1118 studies were reviewed, and 3026 hips were labeled via bounding boxes and classified as normal, displaced femoral neck fracture, nondisplaced femoral neck fracture, intertrochanteric fracture, previous open reduction and internal fixation, or previous arthroplasty. A deep learningโbased object detection model was trained to automate the placement of the bounding boxes. A Densely Connected Convolutional Neural Network (or DenseNet) was trained on a subset of the bounding box images, and its performance was evaluated on a held-out test set and by comparison on a 100-image subset with two groups of human observers: fellowship-trained radiologists and orthopedists; senior residents in emergency medicine, radiology, and orthopedics. The binary accuracy for detecting a fracture of this model was 93.7% (95% confidence interval [CI]: 90.8%, 96.5%), with a sensitivity of 93.2% (95% CI: 88.9%, 97.1%) and a specificity of 94.2% (95% CI: 89.7%, 98.4%).
MIT moves toward greener, more sustainable artificial intelligence
While current artificial intelligence (AI) technology holds strategic and transformative potential, it isn't always environmentally-friendly due to high energy consumption. To the rescue are researchers from Massachusetts Institute of Technology (MIT), who have devised a solution that not only lowers costs but, more importantly, reduces the AI model training's carbon footprint. Back in June 2019, the University of Massachusetts at Amherst revealed that the amount of energy utilized in AI model training equaled 626,000 pounds of carbon dioxide. Contemporary AI isn't just run on a personal laptop or simple server. Rather, deep neural networks are deployed on diverse arrays of specialized hardware platforms. The level of energy consumption required to power such AI technologies is approximately five times the lifetime carbon emissions from an average American car, including its manufacturing.
Hybrid-DNNs: Hybrid Deep Neural Networks for Mixed Inputs
Yuan, Zhenyu, Jiang, Yuxin, Li, Jingjing, Huang, Handong
Rapid development of big data and high-performance computing have encouraged explosive studies of deep learning in geoscience. However, most studies only take single-type data as input, frittering away invaluable multisource, multi-scale information. We develop a general architecture of hybrid deep neural networks (HDNNs) to support mixed inputs. Regarding as a combination of feature learning and target learning, the new proposed networks provide great capacity in high-hierarchy feature extraction and in-depth data mining. Furthermore, the hybrid architecture is an aggregation of multiple networks, demonstrating good flexibility and wide applicability. The configuration of multiple networks depends on application tasks and varies with inputs and targets. Concentrating on reservoir production prediction, a specific HDNN model is configured and applied to an oil development block. Considering their contributions to hydrocarbon production, core photos, logging images and curves, geologic and engineering parameters can all be taken as inputs. After preprocessing, the mixed inputs are prepared as regular-sampled structural and numerical data. For feature learning, convolutional neural networks (CNN) and multilayer perceptron (MLP) network are configured to separately process structural and numerical inputs. Learned features are then concatenated and fed to subsequent networks for target learning. Comparison with typical MLP model and CNN model highlights the superiority of proposed HDNN model with high accuracy and good generalization.
Deep Learning and Bayesian Deep Learning Based Gender Prediction in Multi-Scale Brain Functional Connectivity
Zhao, Gengyan, Hwang, Gyujoon, Cook, Cole J., Liu, Fang, Meyerand, Mary E., Birn, Rasmus M.
Brain gender differences have been known for a long time and are the possible reason for many psychological, psychiatric and behavioral differences between males and females. Predicting genders from brain functional connectivity (FC) can build the relationship between brain activities and gender, and extracting important gender related FC features from the prediction model offers a way to investigate the brain gender difference. Current predictive models applied to gender prediction demonstrate good accuracies, but usually extract individual functional connections instead of connectivity patterns in the whole connectivity matrix as features. In addition, current models often omit the effect of the input brain FC scale on prediction and cannot give any model uncertainty information. Hence, in this study we propose to predict gender from multiple scales of brain FC with deep learning, which can extract full FC patterns as features. We further develop the understanding of the feature extraction mechanism in deep neural network (DNN) and propose a DNN feature ranking method to extract the highly important features based on their contributions to the prediction. Moreover, we apply Bayesian deep learning to the brain FC gender prediction, which as a probabilistic model can not only make accurate predictions but also generate model uncertainty for each prediction. Experiments were done on the high-quality Human Connectome Project S1200 release dataset comprising the resting state functional MRI data of 1003 healthy adults. First, DNN reaches 83.0%, 87.6%, 92.0%, 93.5% and 94.1% accuracies respectively with the FC input derived from 25, 50, 100, 200, 300 independent component analysis (ICA) components. DNN outperforms the conventional machine learning methods on the 25-ICA-component scale FC, but the linear machine learning method catches up as the number of ICA components increases...
ClovaCall: Korean Goal-Oriented Dialog Speech Corpus for Automatic Speech Recognition of Contact Centers
Ha, Jung-Woo, Nam, Kihyun, Kang, Jingu, Lee, Sang-Woo, Yang, Sohee, Jung, Hyunhoon, Kim, Eunmi, Kim, Hyeji, Kim, Soojin, Kim, Hyun Ah, Doh, Kyoungtae, Lee, Chan Kyu, Sung, Nako, Kim, Sunghun
Despite the advancement of ASR, however, most publicly trained from these speech data generally show poor recognition available call-based speech corpora such as Switchboard performance when applied to domain-specific tasks due to the are old-fashioned. Also, most existing call corpora are in English differences in their data distribution and vocabularies. In particular, and mainly focus on open domain dialog or general scenarios AICC requires an accurate ASR model to ensure the precise such as audiobooks. Here we introduce a new large-scale intent classification or slot extraction [9] from user natural Korean call-based speech corpus under a goal-oriented dialog language utterances.
Parsimonious Computing: A Minority Training Regime for Effective Prediction in Large Microarray Expression Data Sets
Sridhar, Shailesh, Saha, Snehanshu, Shaikh, Azhar, Yedida, Rahul, Saha, Sriparna
Rigorous mathematical investigation of learning rates used in back-propagation in shallow neural networks has become a necessity. This is because experimental evidence needs to be endorsed by a theoretical background. Such theory may be helpful in reducing the volume of experimental effort to accomplish desired results. We leveraged the functional property of Mean Square Error, which is Lipschitz continuous to compute learning rate in shallow neural networks. We claim that our approach reduces tuning efforts, especially when a significant corpus of data has to be handled. We achieve remarkable improvement in saving computational cost while surpassing prediction accuracy reported in literature. The learning rate, proposed here, is the inverse of the Lipschitz constant. The work results in a novel method for carrying out gene expression inference on large microarray data sets with a shallow architecture constrained by limited computing resources. A combination of random sub-sampling of the dataset, an adaptive Lipschitz constant inspired learning rate and a new activation function, A-ReLU helped accomplish the results reported in the paper.
Toward Adversarial Robustness by Diversity in an Ensemble of Specialized Deep Neural Networks
Abbasi, Mahdieh, Rajabi, Arezoo, Gagne, Christian, Bobba, Rakesh B.
We aim at demonstrating the influence of diversity in the ensemble of CNNs on the detection of black-box adversarial instances and hardening the generation of white-box adversarial attacks. To this end, we propose an ensemble of diverse specialized CNNs along with a simple voting mechanism. The diversity in this ensemble creates a gap between the predictive confidences of adversaries and those of clean samples, making adversaries detectable. We then analyze how diversity in such an ensemble of specialists may mitigate the risk of the black-box and white-box adversarial examples. Using MNIST and CIFAR-10, we empirically verify the ability of our ensemble to detect a large portion of well-known black-box adversarial examples, which leads to a significant reduction in the risk rate of adversaries, at the expense of a small increase in the risk rate of clean samples. Moreover, we show that the success rate of generating white-box attacks by our ensemble is remarkably decreased compared to a vanilla CNN and an ensemble of vanilla CNNs, highlighting the beneficial role of diversity in the ensemble for developing more robust models.
The critical locus of overparameterized neural networks
Many aspects of the geometry of loss functions in deep learning remain mysterious. In this paper, we work toward a better understanding of the geometry of the loss function $L$ of overparameterized feedforward neural networks. In this setting, we identify several components of the critical locus of $L$ and study their geometric properties. For networks of depth $\ell \geq 4$, we identify a locus of critical points we call the star locus $S$. Within $S$ we identify a positive-dimensional sublocus $C$ with the property that for $p \in C$, $p$ is a degenerate critical point, and no existing theoretical result guarantees that gradient descent will not converge to $p$. For very wide networks, we build on earlier work and show that all critical points of $L$ are degenerate, and give lower bounds on the number of zero eigenvalues of the Hessian at each critical point. For networks that are both deep and very wide, we compare the growth rates of the zero eigenspaces of the Hessian at all the different families of critical points that we identify. The results in this paper provide a starting point to a more quantitative understanding of the properties of various components of the critical locus of $L$.
Deep Learning Based Integrators for Solving Newton's Equations with Large Timesteps
Kadupitiya, JCS, Fox, Geoffrey C., Jadhao, Vikram
Classical molecular dynamics simulations are based on Newton's equations of motion and rely on numerical integrators to solve them. Using a small timestep to avoid discretization errors, Verlet integrators generate a trajectory of particle positions as solutions to Newton's equations. We introduce an integrator based on deep neural networks that is trained on trajectories generated using the Verlet integrator and learns to propagate the dynamics of particles with timestep up to 4000$\times$ larger compared to the Verlet timestep. We demonstrate significant net speedup of up to 32000 for 1 - 16 particle 3D systems and over a variety of force fields.