Goto

Collaborating Authors

 Country


Improving Neural Story Generation by Targeted Common Sense Grounding

arXiv.org Machine Learning

Stories generated with neural language models have shown promise in grammatical and stylistic consistency. However, the generated stories are still lacking in common sense reasoning, e.g., they often contain sentences deprived of world knowledge. W e propose a simple multi-task learning scheme to achieve quantitatively better common sense reasoning in language models by leveraging auxiliary training signals from datasets designed to provide common sense grounding. When combined with our two-stage fine-tuning pipeline, our method achieves improved common sense reasoning and state-of-the-art perplexity on the Writing-Prompts ( Fan et al., 2018) story generation dataset.


Theoretical Issues in Deep Networks: Approximation, Optimization and Generalization

arXiv.org Machine Learning

While deep learning is successful in a number of applications, it is not yet well understood theoretically. A satisfactory theoretical characterization of deep learning however, is beginning to emerge. It covers the following questions: 1) representation power of deep networks 2) optimization of the empirical risk 3) generalization properties of gradient descent techniques --- why the expected error does not suffer, despite the absence of explicit regularization, when the networks are overparametrized? In this review we discuss recent advances in the three areas. In approximation theory both shallow and deep networks have been shown to approximate any continuous functions on a bounded domain at the expense of an exponential number of parameters (exponential in the dimensionality of the function). However, for a subset of compositional functions, deep networks of the convolutional type can have a linear dependence on dimensionality, unlike shallow networks. In optimization we discuss the loss landscape for the exponential loss function and show that stochastic gradient descent will find with high probability the global minima. To address the question of generalization for classification tasks, we use classical uniform convergence results to justify minimizing a surrogate exponential-type loss function under a unit norm constraint on the weight matrix at each layer -- since the interesting variables for classification are the weight directions rather than the weights. Our approach, which is supported by several independent new results, offers a solution to the puzzle about generalization performance of deep overparametrized ReLU networks, uncovering the origin of the underlying hidden complexity control.


Almost Tune-Free Variance Reduction

arXiv.org Machine Learning

The variance reduction class of algorithms including the representative ones, abbreviated as SVRG and SARAH, have well documented merits for empirical risk minimization tasks. However, they require grid search to optimally tune parameters (step size and the number of iterations per inner loop) for best performance. This work introduces `almost tune-free' SVRG and SARAH schemes by equipping them with Barzilai-Borwein (BB) step sizes. To achieve the best performance, both i) averaging schemes; and, ii) the inner loop length are adjusted according to the BB step size. SVRG and SARAH are first reexamined through an `estimate sequence' lens. Such analysis provides new averaging methods that tighten the convergence rates of both SVRG and SARAH theoretically, and improve their performance empirically when the step size is chosen large. Then a simple yet effective means of adjusting the number of iterations per inner loop is developed, which completes the tune-free variance reduction together with BB step sizes. Numerical tests corroborate the proposed methods.


Principal Component Analysis Using Structural Similarity Index for Images

arXiv.org Machine Learning

Despite the advances of deep learning in specific tasks using images, the principled assessment of image fidelity and similarity is still a critical ability to develop. As it has been shown that Mean Squared Error (MSE) is insufficient for this task, other measures have been developed with one of the most effective being Structural Similarity Index (SSIM). Such measures can be used for subspace learning but existing methods in machine learning, such as Principal Component Analysis (PCA), are based on Euclidean distance or MSE and thus cannot properly capture the structural features of images. In this paper, we define an image structure subspace which discriminates different types of image distortions. We propose Image Structural Component Analysis (ISCA) and also kernel ISCA by using SSIM, rather than Euclidean distance, in the formulation of PCA. This paper provides a bridge between image quality assessment and manifold learning opening a broad new area for future research.


Predicting the Long-Term Outcomes of Biologics in Psoriasis Patients Using Machine Learning

arXiv.org Machine Learning

Background. Real-world data show that approximately 50% of psoriasis patients treated with a biologic agent will discontinue the drug because of loss of efficacy. History of previous therapy with another biologic, female sex and obesity were identified as predictors of drug discontinuations, but their individual predictive value is low. Objectives. To determine whether machine learning algorithms can produce models that can accurately predict outcomes of biologic therapy in psoriasis on individual patient level. Results. All tested machine learning algorithms could accurately predict the risk of drug discontinuation and its cause (e.g. lack of efficacy vs adverse event). The learned generalized linear model achieved diagnostic accuracy of 82%, requiring under 2 seconds per patient using the psoriasis patients dataset. Input optimization analysis established a profile of a patient who has best chances of long-term treatment success: biologic-naive patient under 49 years, early-onset plaque psoriasis without psoriatic arthritis, weight < 100 kg, and moderate-to-severe psoriasis activity (DLQI $\geq$ 16; PASI $\geq$ 10). Moreover, a different generalized linear model is used to predict the length of treatment for each patient with mean absolute error (MAE) of 4.5 months. However Pearson Correlation Coefficient indicates 0.935 linear dependencies between the actual treatment lengths and predicted ones. Conclusions. Machine learning algorithms predict the risk of drug discontinuation and treatment duration with accuracy exceeding 80%, based on a small set of predictive variables. This approach can be used as a decision-making tool, communicating expected outcomes to the patient, and development of evidence-based guidelines.


Tutorial and Survey on Probabilistic Graphical Model and Variational Inference in Deep Reinforcement Learning

arXiv.org Artificial Intelligence

Probabilistic Graphical Modeling and Variational Inference play an important role in recent advances in Deep Reinforcement Learning. Aiming at a self-consistent tutorial survey, this article illustrates basic concepts of reinforcement learning with Probabilistic Graphical Models, as well as derivation of some basic formula as a recap. Reviews and comparisons on recent advances in deep reinforcement learning with different research directions are made from various aspects. We offer Probabilistic Graphical Models, detailed explanation and derivation to several use cases of Variational Inference, which serve as a complementary material on top of the original contributions.


Exploring the Performance of Deep Residual Networks in Crazyhouse Chess

arXiv.org Artificial Intelligence

Crazyhouse is a chess variant that incorporates all of the classical chess rules, but allows users to drop pieces captured from the opponent as a normal move. Until 2018, all competitive computer engines for this board game made use of an alpha-beta pruning algorithm with a hand-crafted evaluation function for each position. Previous machine learning-based algorithms for just regular chess, such as NeuroChess and Giraffe, took hand-crafted evaluation features as input rather than a raw board representation. More recent projects, such as AlphaZero, reached massive success but required massive computational resources in order to reach its final strength. This paper describes the development of SixtyFour, an engine designed to compete in the chess variant of Crazyhouse with limited hardware. This specific variant poses a multitude of significant challenges due to its large branching factor, state-space complexity, and the multiple move types a player can make. We propose the novel creation of a neural network-based evaluation function for Crazyhouse. More importantly, we evaluate the effectiveness of an ensemble model, which allows the training time and datasets to be easily distributed on regular CPU hardware commodity. Early versions of the network have attained a playing level comparable to a strong amateur on online servers.


DIU announces disaster-mapping AI challenge - FedScoop

#artificialintelligence

The Defense Innovation Unit is looking for participants in its second computer-vision artificial intelligence challenge, this time to identify building damage in post-disaster areas. The xVIEW2 challenge, to start in early September, asks machine learning experts to develop algorithms and models to analyze post-disaster satellite imagery to improve mapping. Understanding the scale of disasters is still an often dangerous and time-consuming process that slows first responders. With AI analyzing satellite imagery, the Department of Defense hopes to improve response time and effectiveness in those critical moments. Competitors will be judged on their ability to identify buildings and score how badly damaged they are.


Artificial Intelligence (AI) Stats News: 50% Of Americans Optimistic And 50% Fearful About AI

#artificialintelligence

The recent surveys, studies, forecasts and other quantitative assessments of the health and progress of AI found that Americans are evenly divided over its promise or peril while not entirely sure what it is; that they are increasingly unhappy with technology companies that develop AI; that psychiatrists don't see AI replacing them but ad copywriters may want to consider other vocations; that the data that feeds and nourish AI keeps getting into the wrong hands; and that AI may assist radiologists, cardiologists, and fast-food chains. Glass half full? 50% of American consumers feel "optimistic and informed" about AI while the other half feel "fearful and uninformed" about AI [Blumberg Capital surveys of 1,000 U.S. consumers aged 18 ] Do you trust AI? 67% of Americans believe that self-driving cars will be safer than human-operated cars; 44% say that if a self-driving Uber car picked them up, they would get in; 87% say a licensed driver should be behind the wheel ready to take control if needed; 35% say they would never drive in a self-driving car [DriversED.com Do you trust the companies that develop AI? Only 50% of Americans believe technology companies have a positive impact on their country, down from 71% four years ago. Negative views of technology companies' impact on the U.S. have nearly doubled during this period, from 17% to 33% [Pew Research Center phone survey of 1,502 adults July 2019] AI may not replace psychiatrists: Only 3.8% of 791 psychiatrists surveyed in 22 countries felt that AI/ML was likely to replace a human clinician for providing empathetic care; documenting (e.g.


AutoML GAN AutoGAN! AI Can Now Design Better GAN Models Than Humans

#artificialintelligence

Thanks to the creation of AutoML -- which is essentially automated neural architecture search (NAS) -- AI can now design better deep neural networks than human researchers for computer vision tasks such as image classification and object detection. AutoML's tremendous success has prompted AI researchers to explore its efficacy in additional areas, such as generative adversarial networks (GANs). Researchers from Texas A&M University and MIT-IBM Watson AI Lab recently presented a paper that applies NAS to GANs. Their "AutoGAN" is an architecture search scheme specifically tailored for GANs that outperforms current state-of-the-art hand-crafted GANs on the task of unconditional image generation. The associated paper has been accepted by ICCV 2019.