Country
My Computer Is an Honor Student — but How Intelligent Is It? Standardized Tests as a Measure of AI
Clark, Peter (Allen Institute for Artificial Intelligence) | Etzioni, Oren (Allen Institute for Artificial Intelligence)
Given the well-known limitations of the Turing Test, there is a need for objective tests to both focus attention on, and measure progress towards, the goals of AI. In this paper we argue that machine performance on standardized tests should be a key component of any new measure of AI, because attaining a high level of performance requires solving significant AI problems involving language understanding and world modeling - critical skills for any machine that lays claim to intelligence. In addition, standardized tests have all the basic requirements of a practical test: they are accessible, easily comprehensible, clearly measurable, and offer a graduated progression from simple tasks to those requiring deep understanding of the world. Here we propose this task as a challenge problem for the community, summarize our state-of-the-art results on math and science tests, and provide supporting datasets
Beyond the Turing Test
Marcus, Gary (New York University) | Rossi, Francesca (University of Padova) | Veloso, Manuela (Carnegie Mellon University)
Within the field, the test is widely recognized as a pioneering landmark, but also is now seen as a distraction, designed over half a century ago, and too crude to really measure intelligence. Intelligence is, after all, a multidimensional variable, and no one test could possibly ever be definitive truly to measure it. Moreover, the original test, at least in its standard implementations, has turned out to be highly gameable, arguably an exercise in deception rather than a true measure of anything especially correlated with intelligence. The much ballyhooed 2015 Turing test winner Eugene Goostman, for instance, pretends to be a thirteen-year-old foreigner and proceeds mainly by ducking questions and returning canned one-liners; it cannot see, it cannot think, and it is certainly a long way from genuine artificial general intelligence.
Inverse Reinforcement Learning with Simultaneous Estimation of Rewards and Dynamics
Herman, Michael, Gindele, Tobias, Wagner, Jörg, Schmitt, Felix, Burgard, Wolfram
Inverse Reinforcement Learning (IRL) describes the problem of learning an unknown reward function of a Markov Decision Process (MDP) from observed behavior of an agent. Since the agent's behavior originates in its policy and MDP policies depend on both the stochastic system dynamics as well as the reward function, the solution of the inverse problem is significantly influenced by both. Current IRL approaches assume that if the transition model is unknown, additional samples from the system's dynamics are accessible, or the observed behavior provides enough samples of the system's dynamics to solve the inverse problem accurately. These assumptions are often not satisfied. To overcome this, we present a gradient-based IRL approach that simultaneously estimates the system's dynamics. By solving the combined optimization problem, our approach takes into account the bias of the demonstrations, which stems from the generating policy. The evaluation on a synthetic MDP and a transfer learning task shows improvements regarding the sample efficiency as well as the accuracy of the estimated reward functions and transition models.
A Differentiable Transition Between Additive and Multiplicative Neurons
Köpp, Wiebke, van der Smagt, Patrick, Urban, Sebastian
A BSTRACT Existing approaches to combine both additive and multiplicative neural units either use a fixed assignment of operations or require discrete optimization to determine what function a neuron should perform. However, this leads to an extensive increase in the computational complexity of the training procedure. We present a novel, parameterizable transfer function based on the mathematical concept of non-integer functional iteration that allows the operation each neuron performs to be smoothly and, most importantly, differentiablely adjusted between addition and multiplication. This allows the decision between addition and multiplication to be integrated into the standard backpropagation training procedure. The value of such a product unit is given byy i σ ( j x W ij j).
Loss Functions for Top-k Error: Analysis and Insights
Lapin, Maksim, Hein, Matthias, Schiele, Bernt
In order to push the performance on realistic computer vision tasks, the number of classes in modern benchmark datasets has significantly increased in recent years. This increase in the number of classes comes along with increased ambiguity between the class labels, raising the question if top-1 error is the right performance measure. In this paper, we provide an extensive comparison and evaluation of established multiclass methods comparing their top-k performance both from a practical as well as from a theoretical perspective. Moreover, we introduce novel top-k loss functions as modifications of the softmax and the multiclass SVM losses and provide efficient optimization schemes for them. In the experiments, we compare on various datasets all of the proposed and established methods for top-k error optimization. An interesting insight of this paper is that the softmax loss yields competitive top-k performance for all k simultaneously. For a specific top-k error, our new top-k losses lead typically to further improvements while being faster to train than the softmax.
A Linearly-Convergent Stochastic L-BFGS Algorithm
Moritz, Philipp, Nishihara, Robert, Jordan, Michael I.
We propose a new stochastic L-BFGS algorithm and prove a linear convergence rate for strongly convex and smooth functions. Our algorithm draws heavily from a recent stochastic variant of L-BFGS proposed in Byrd et al. (2014) as well as a recent approach to variance reduction for stochastic gradient descent from Johnson and Zhang (2013). We demonstrate experimentally that our algorithm performs well on large-scale convex and non-convex optimization problems, exhibiting linear convergence and rapidly solving the optimization problems to high levels of precision. Furthermore, we show that our algorithm performs well for a wide-range of step sizes, often differing by several orders of magnitude.
Bayesian inference in hierarchical models by combining independent posteriors
Dutta, Ritabrata, Blomstedt, Paul, Kaski, Samuel
Noname manuscript No. (will be inserted by the editor) Abstract Hierarchical models are versatile tools for joint modeling of data sets arising from different, but related, sources. Fully Bayesian inference may, however, become computationally prohibitive if the sourcespecific data models are complex, or if the number of sources is very large. To facilitate computation, we propose an approach, where inference is first made independently for the parameters of each data set, whereupon the obtained posterior samples are used as observed data in a substitute hierarchical model, based on a scaled likelihood function. Compared to direct inference in a full hierarchical model, the approach has the advantage of being able to speed up convergenceby breaking down the initial large inference problem into smaller individual subproblems with better convergence properties. Moreover it enables parallel processing of the possibly complex inferences of the source-specific parameters, which may otherwise create a computational bottleneck if processed jointly as part of a hierarchical model.
The 5-Minute Interview: Rachel Lader, Software Engineer, Apple
Keywords: 5-minute interview apple cypher express.js Bryce Merkl Sasaki is the Community Content Manager for Neo Technology. He studied professional and creative writing for undergrad and has been freelancing for 7 years. Recently, he worked at an inbound marketing agency in Philadelphia as a copywriter before moving to California. When not working, he likes to spend his time working on his novel, looking for pickup soccer games and reading voraciously.
Accurate Sales Forecast for Data Analysts: Building a Random Forest model with Just SQL and Hivemall Treasure Data Blog
In this blog post, we will use Hivemall, the open source Machine Learning-on-SQL library available in the Treasure Data environment, to introduce the basics of machine learning. We will use an E-Commerce dataset from Kaggle, the data science competition platform. The first challenge is predicting the retail sales for the Rossman stores (the full details at Kaggle). We will use an ensemble learning technique known as Random Forest regression. Rossman is a pharmacy chain with over 3,000 stores in seven countries within Europe.
DataRPM & Tamr partner to deliver end-to-end automation of machine learning from data ingestion to strategic business insights for their customers
WIRE)--DataRPM, the award-winning Cognitive Data Science company that automates Machine Learning to deliver Recommendation & Prediction data products, today announced a partnership with big data analytics company, Tamr, Inc. "We are truly excited about our partnership with Tamr to expand the portfolio of innovative solutions to provide end-to-end data products," said Sundeep Sanghavi, DataRPM Co-Founder and Chief Executive Officer. "Data integration quality and prep is a crucial requirement for every big data initiative. Our partnership further meets the expanding needs of our customers and allows them to truly leverage machine learning all the way from data ingestion to strategic business insights." "DataRPM's Cognitive Data Science Platform automates machine learning for Recommendations & Predictions to deliver continuous insights that propel enterprises' growth dramatically, thus redefining data science with scale, speed and repeatability," further explained Sundeep. "With the dramatic explosion in data, compounded by the growing shortage of data scientists, and the need for faster sprint cycles to launch strategic initiatives, the call of the hour is Cognitive Data Science. By automating Machine Learning, greater value from productized data can now be derived through operationalizing its usage within companies' process flows in a continuous manner. "Tamr's machine-driven, human-guided approach to data preparation aligns closely with DataRPM's machine learning automation," said Andy Palmer, Tamr Co-Founder and Chief Executive Officer. "DataRPM's Prediction and Recommendation models need a continuous flow of clean, unified data that Tamr provides.