Genre
Nested Mini-Batch K-Means
Newling, James, Fleuret, François
A new algorithm is proposed which accelerates the mini-batch k-means algorithm of Sculley (2010) by using the distance bounding approach of Elkan (2003). We argue that, when incorporating distance bounds into a mini-batch algorithm, already used data should preferentially be reused. To this end we propose using nested mini-batches, whereby data in a mini-batch at iteration t is automatically reused at iteration t+1. Using nested mini-batches presents two difficulties. The first is that unbalanced use of data can bias estimates, which we resolve by ensuring that each data sample contributes exactly once to centroids. The second is in choosing mini-batch sizes, which we address by balancing premature fine-tuning of centroids with redundancy induced slow-down. Experiments show that the resulting nmbatch algorithm is very effective, often arriving within 1% of the empirical minimum 100 times earlier than the standard mini-batch algorithm.
Adaptive matching pursuit for sparse signal recovery
Vu, Tiep H., Mousavi, Hojjat S., Monga, Vishal
Spike and Slab priors have been of much recent interest in signal processing as a means of inducing sparsity in Bayesian inference. Applications domains that benefit from the use of these priors include sparse recovery, regression and classification. It is well-known that solving for the sparse coefficient vector to maximize these priors results in a hard non-convex and mixed integer programming problem. Most existing solutions to this optimization problem either involve simplifying assumptions/relaxations or are computationally expensive. We propose a new greedy and adaptive matching pursuit (AMP) algorithm to directly solve this hard problem. Essentially, in each step of the algorithm, the set of active elements would be updated by either adding or removing one index, whichever results in better improvement. In addition, the intermediate steps of the algorithm are calculated via an inexpensive Cholesky decomposition which makes the algorithm much faster. Results on simulated data sets as well as real-world image recovery challenges confirm the benefits of the proposed AMP, particularly in providing a superior cost-quality trade-off over existing alternatives.
Policy Networks with Two-Stage Training for Dialogue Systems
Fatemi, Mehdi, Asri, Layla El, Schulz, Hannes, He, Jing, Suleman, Kaheer
In this paper, we propose to use deep policy networks which are trained with an advantage actor-critic method for statistically optimised dialogue systems. First, we show that, on summary state and action spaces, deep Reinforcement Learning (RL) outperforms Gaussian Processes methods. Summary state and action spaces lead to good performance but require pre-engineering effort, RL knowledge, and domain expertise. In order to remove the need to define such summary spaces, we show that deep RL can also be trained efficiently on the original state and action spaces. Dialogue systems based on partially observable Markov decision processes are known to require many dialogues to train, which makes them unappealing for practical deployment. We show that a deep RL method based on an actor-critic architecture can exploit a small amount of data very efficiently. Indeed, with only a few hundred dialogues collected with a handcrafted policy, the actor-critic deep learner is considerably boot-strapped from a combination of supervised and batch RL. In addition, convergence to an optimal policy is significantly sped up compared to other deep RL methods initialized on the data with batch RL. All experiments are performed on a restaurant domain derived from the Dialogue State Tracking Challenge 2 (DSTC2) dataset.
Modelling Creativity: Identifying Key Components through a Corpus-Based Approach
As Torrance observes: '[c]reativity defies precise definition... even if we had a precise conception of creativity, I am certain we would have difficulty putting it into words' [15, p. 43]. Many other authors have expressed similar difficulties [7, 10, 16]. In their review of research into human creativity, Hennessey and Amabile ask a significant follow-on question: 'Even if this mysterious phenomenon can be isolated, quantified, and dissected, why bother? Wouldn't it make more sense to revel in the mystery and wonder of it all?' [11, p. 570] Two answers to this question are offered by Hennessey and Amabile, both of which are identified as desirable: to gain a deeper understanding of creativity and to learn how to boost people's creativity. Creativity can and should be studied and measured scientifically, but the lack of a commonly-agreed understanding causes problems for measurement [10]. Plucker et al. make recommendations about best practice based on their own survey of the creativity literature: 'we argue that creativity researchers must (a) explicitly define what they mean by creativity, (b) avoid using scores of creativity measures as the sole definition of creativity (e.g., creativity is what creativity tests measure and creativity tests measure creativity, therefore we will use a score on a creativity test as our outcome variable), (c) discuss how the definition they are using is similar to or different from other definitions, and (d) address the question of creativity for whom and in what context.' [9, p.92] In short, we need to specify and justify the standards that we use to judge creativity. A more objective and well-articulated account of how creativity is manifested enables researchers to make a worthwhile contribution [8-10]. Particularly, in research we would like to focus on what processes and concepts relevant to creativity are'sufficiently important to warrant study' [17, p. 15], based on an accumulation of the body of work on creativity to date [17].
Meet Your Robot Pharmacist
But while the San Francisco entrepreneur misses his mother and father in Australia, he doesn't worry about their health. That's because he's pinged multiple times per day about their medication management and activity levels via an app connected to the PillDrill Wi-Fi health hub in their home. When the parents pop a pill, the system updates that the medication has been ingested and even registers their current state of mind with a scannable mood cube -- a different emotion is pictured on five of the six sides to represent comfort levels and pain. Havas, the founder of the company that manufactures the PillDrill, isn't trying to be creepy; he's trying to be cognizant of his parents' health care -- and stay ahead of trouble. In 2012, around 300,000 Americans called poison control hotlines after accidentally ingesting medications, according to the National Poison Control Center.
Machine Learning in a Year – Learning New Stuff
During the christmas vacation of 2015, I got a motivational boost again and decided try out Kaggle. So I spent quite some time experimenting with various algorithms for their Homesite Quote Conversion, Otto Group Product Classification and Bike Sharing Demand contests. The main takeaway from this was the experience of iteratively improving the results by experimenting with the algorithms and the data. I learned to trust my logic when doing machine learning. If tweaking a parameter or engineering a new feature seems like a good idea logically, it's quite likely that it actually will help.
nick lally // art, geography, software » Blog Archive » geographies of software, AAG 2017
A variety of technologies have emerged in the last decade that make it easier and cheaper than ever before to make representations of everyday mobile embodiment. Increasing numbers of people are quantifying and self-tracking their everyday lives recording behavioural, biological and environmental data (Beer, 2016; Neff & Nafus, 2016) using a variety of technologies, for example: • lightweight wearable cameras such as the GoPro allowing users to record footage of their most banal everyday activities; • devices such as the Fitbit and Apple Watch bringing continuous physiological monitoring out of the medical realm and into mainstream culture; • apps like Strava allowing people to quantify their cycling, running and walking activities; • lightweight devices for measuring brain activity (EEG) and stimulation (EDA) becoming sufficiently robust and discreet to be used in non-lab environments. None of the underlying technologies are novel, but as they are made accessible in cheaper and more user-friendly packages, new techniques and sources of data are becoming more readily available for geographical analysis. Engagement with these technologies has created a rapidly expanding area of investigation within geography. The emergence of the quantified-self poses both opportunities and dilemmas for geographical thought. We wish to move past simplistic protests that dismiss such technology as offering another take on Haraway's (1988) 'god trick', presenting partial, and highly situated data as objective truth. Instead, this session will build on the potential identified by Delyser and Sui (2013) to take more inventive approaches toward mobile methods. The focus will be on how these technologies can be engaged with by critical geographers to bring new perspectives to their analysis of everyday embodiment.
A 3D-printed autonomous car, and more in the week that was
That's the idea behind Local Motors' latest vehicle, which features a 3D-printed body, a windshield video screen and no steering wheel. Meanwhile, OX launched the world's first all-terrain flat-pack truck, which can be quickly shipped anywhere in the world. Cannae Corporation announced plans to test an "impossible" zero-exhaust microwave thruster that could revolutionize space travel. And Electra Meccanica launched SOLO, an affordable three-wheeled electric vehicle for one. In energy news, this week Sonos Motors announced plans to debut a solar-powered car within two years, and Soel Yachts unveiled a sun-powered motorboat that glides through the water without making a sound.
Book Review: Python Machine Learning by Sebastian Raschka
Machine learning and AI seem to be the way of the future. These techniques can help software developers create powerful applications that crunch data, analyze trends, and offer solutions that a developer may not even consider. But getting into machine learning is no easy task. You need to have a background in programming and your algorithm skills need to be at least somewhat competent. Python Machine Learning is one detailed book that covers machine learning from the angle of 3rd party Python libraries.
Peircean Induction and the Error-Correcting Thesis (Part I)
Today is C.S. Peirce's birthday. You should read him: he's a treasure chest on essentially any topic, and he anticipated several major ideas in statistics (e.g., randomization, confidence intervals) as well as in logic. Links to Parts 2 and 3 are at the end. It's written for a very general philosophical audience; the statistical parts are pretty informal. Peirce's philosophy of inductive inference in science is based on the idea that what permits us to make progress in science, what allows our knowledge to grow, is the fact that science uses methods that are self-correcting or error-correcting: Induction is the experimental testing of a theory.