Deep Learning
Evaluating Lossy Compression Rates of Deep Generative Models
Huang, Sicong, Makhzani, Alireza, Cao, Yanshuai, Grosse, Roger
The field of deep generative modeling has succeeded in producing astonishingly realistic-seeming images and audio, but quantitative evaluation remains a challenge. Log-likelihood is an appealing metric due to its grounding in statistics and information theory, but it can be challenging to estimate for implicit generative models, and scalar-valued metrics give an incomplete picture of a model's quality. In this work, we propose to use rate distortion (RD) curves to evaluate and compare deep generative models. While estimating RD curves is seemingly even more computationally demanding than log-likelihood estimation, we show that we can approximate the entire RD curve using nearly the same computations as were previously used to achieve a single log-likelihood estimate. We evaluate lossy compression rates of VAEs, GANs, and adversarial autoencoders (AAEs) on the MNIST and CIFAR10 datasets. Measuring the entire RD curve gives a more complete picture than scalar-valued metrics, and we arrive at a number of insights not obtainable from log-likelihoods alone.
Peer-inspired Student Performance Prediction in Interactive Online Question Pools with Graph Neural Network
Li, Haotian, Wei, Huan, Wang, Yong, Song, Yangqiu, Qu, Huamin
Student performance prediction is critical to online education. It can benefit many downstream tasks on online learning platforms, such as estimating dropout rates, facilitating strategic intervention, and enabling adaptive online learning. Interactive online question pools provide students with interesting interactive questions to practice their knowledge in online education. However, little research has been done on student performance prediction in interactive online question pools. Existing work on student performance prediction targets at online learning platforms with predefined course curriculum and accurate knowledge labels like MOOC platforms, but they are not able to fully model knowledge evolution of students in interactive online question pools. In this paper, we propose a novel approach using Graph Neural Networks (GNNs) to achieve better student performance prediction in interactive online question pools. Specifically, we model the relationship between students and questions using student interactions to construct the student-interaction-question network and further present a new GNN model, called R^2GCN, which intrinsically works for the heterogeneous networks, to achieve generalizable student performance prediction in interactive online question pools. We evaluate the effectiveness of our approach on a real-world dataset consisting of 104,113 mouse trajectories generated in the problem-solving process of over 4000 students on 1631 questions. The experiment results show that our approach can achieve a much higher accuracy of student performance prediction than both traditional machine learning approaches and GNN models.
Simple and Effective VAE Training with Calibrated Decoders
Rybkin, Oleh, Daniilidis, Kostas, Levine, Sergey
Variational autoencoders (VAEs) provide an effective and simple method for modeling complex distributions. However, training VAEs often requires considerable hyperparameter tuning, and often utilizes a heuristic weight on the prior KL-divergence term. In this work, we study how the performance of VAEs can be improved while not requiring the use of this heuristic hyperparameter, by learning calibrated decoders that accurately model the decoding distribution. While in some sense it may seem obvious that calibrated decoders should perform better than uncalibrated decoders, much of the recent literature that employs VAEs uses uncalibrated Gaussian decoders with constant variance. We observe empirically that the na\"{i}ve way of learning variance in Gaussian decoders does not lead to good results. However, other calibrated decoders, such as discrete decoders or learning shared variance can substantially improve performance. To further improve results, we propose a simple but novel modification to the commonly used Gaussian decoder, which represents the prediction variance non-parametrically. We observe empirically that using the heuristic weight hyperparameter is not necessary with our method. We analyze the performance of various discrete and continuous decoders on a range of datasets and several single-image and sequential VAE models. Project website: https://orybkin.github.io/sigma-vae/
Chrome Dino Run using Reinforcement Learning
Marwah, Divyanshu, Srivastava, Sneha, Gupta, Anusha, Verma, Shruti
Reinforcement Learning is one of the most advanced set of algorithms known to mankind which can compete in games and perform at par or even better than humans. In this paper we study most popular model free reinforcement learning algorithms along with convolutional neural network to train the agent for playing the game of Chrome Dino Run. We have used two of the popular temporal difference approaches namely Deep Q-Learning, and Expected SARSA and also implemented Double DQN model to train the agent and finally compare the scores with respect to the episodes and convergence of algorithms with respect to timesteps.
Neural Networks from Scratch with Python Code and Math in Detail-- I
Note: In our second tutorial on neural networks, we dive in-depth on the limitations and advantages of using neural networks. We show how to implement neural nets with hidden layers and how these lead to a higher accuracy rate on our predictions, along with implementation samples in Python on Google Colab. Neural networks form the base of deep learning, which is a subfield of machine learning, where the structure of the human brain inspires the algorithms. Neural networks take input data, train themselves to recognize patterns found in the data, and then predict the output for a new set of similar data. Therefore, a neural network can be thought of as the functional unit of deep learning, which mimics the behavior of the human brain to solve complex data-driven problems.
3 Common Challenges That Deep Learning Faces In Medical Imaging
Medical Imaging is one of the popular fields where the researchers are widely exploring deep learning. But the research may not translate easily into a practical or production-ready tech. In an engaging session by Abdul Jilani at the Computer Vision Developer Conference 2020, Abdul Jilani who is the lead data scientist at DataRobot explained the various challenges that applied machine learning face and how these can be overcome in real-time environments. "Building good training datasets and the wrong practices which lead to data leakage were the most commonly faced challenges that I investigated in a famous COVID X-rays prediction paper," he said. He further shared examples where Google's medical AI was super accurate in the lab, but the real-life story turned out to be very different, as it failed to give results at all when applied to real-life situations.
Deep neural networks enable quantitative movement analysis using single-camera videos
Many neurological and musculoskeletal diseases impair movement, which limits peopleโs function and social participation. Quantitative assessment of motion is critical to medical decision-making but is currently possible only with expensive motion capture systems and highly trained personnel. Here, we present a method for predicting clinically relevant motion parameters from an ordinary video of a patient. Our machine learning models predict parameters include walking speed (rโ=โ0.73), cadence (rโ=โ0.79), knee flexion angle at maximum extension (rโ=โ0.83), and Gait Deviation Index (GDI), a comprehensive metric of gait impairment (rโ=โ0.75). These correlation values approach the theoretical limits for accuracy imposed by natural variability in these metrics within our patient population. Our methods for quantifying gait pathology with commodity cameras increase access to quantitative motion analysis in clinics and at home and enable researchers to conduct large-scale studies of neurological and musculoskeletal disorders. In the context of diseases impairing movement, quantitative assessment of motion is critical to medical decision-making but is currently possible only with expensive motion capture systems and trained personnel. Here, the authors present a method for predicting clinically relevant motion parameters from an ordinary video of a patient.
A college kid used AI to create a fake blog. It reached #1 on Hacker News.
GPT-3 is OpenAI's latest and largest language AI model, which the San Franciscoโbased research lab began drip-feeding out in mid-July. In February of last year, OpenAI made headlines with GPT-2, an earlier version of the algorithm, which it announced it would withhold for fear it would be abused. The decision immediately sparked a backlash, as researchers accused the lab of pulling a stunt. By November, the lab had reversed position and released the model, saying it had detected "no strong evidence of misuse so far." The lab took a different approach with GPT-3; it neither withheld it nor granted public access.
Article - Artificial Intelligence and Imaging Optimization
In recent years, through its ability to collect and swiftly analyze huge volumes of data generated by imaging studies, artificial intelligence (AI) has been revolutionizing the practice of radiology. Throughout the field, applications leveraging AI are being used to improve diagnostic accuracy, imaging consistency, workflow efficiency, and patient care by automating many formerly tedious, time-consuming, and manually performed tasks. Traditional methods for implementing AI have relied primarily on machine-learning algorithms based on expert programming of predefined rules. More recent advances, however, have given rise to superior algorithms that learn through direct navigation of data to "recognize" potentially suspicious findings. These "deep learning" algorithms are advantageous because they operate with minimal human expert intervention, and instead collect and process data in raw form through the use of artificial neural networks.
GPT-3, explained: This new language AI is uncanny, funny -- and a big deal
Last month, OpenAI, the Elon Musk-founded artificial intelligence research lab, announced the arrival of the newest version of an AI system it had been working on that can mimic human language, a model called GPT-3. In the weeks that followed, people got the chance to play with the program. If you follow news about AI, you may have seen some headlines calling it a huge step forward, even a scary one. I've now spent the past few days looking at GPT-3 in greater depth and playing around with it. I'm here to tell you: The hype is real. It has its shortcomings, but make no mistake: GPT-3 represents a tremendous leap for AI. A year ago I sat down to play with GPT-3's precursor dubbed (you guessed it) GPT-2.