Goto

Collaborating Authors

 Education


Stochastic Cubic Regularization for Fast Nonconvex Optimization

arXiv.org Machine Learning

In this setting, we only have access to the stochastic function f(x; ξ), where the random variable ξ is sampled from an underlying distribution D. The task is to optimize the expected function f(x), which in general may be nonconvex. This framework covers a wide range of problems, including the offline setting where we minimize the empirical loss over a fixed amount of data, and the online setting where data arrives sequentially. One of the most prominent applications of stochastic optimization has been in large-scale statistics and machine learning problems, such as the optimization of deep neural networks. Classical analysis in nonconvex optimization only guarantees convergence to a first-order stationary point (i.e., a point x satisfying ‖ f(x)‖ 0), which can be a local minimum, a local maximum, or a saddle point. This paper goes further, proposing an algorithm that escapes saddle points and converges to a local minimum.


On consistent vertex nomination schemes

arXiv.org Machine Learning

Given a vertex of interest in a network $G_1$, the vertex nomination problem seeks to find the corresponding vertex of interest (if it exists) in a second network $G_2$. Although the vertex nomination problem and related tasks have attracted much attention in the machine learning literature, with applications to social and biological networks, the framework has so far been confined to a comparatively small class of network models, and the concept of statistically consistent vertex nomination schemes has been only shallowly explored. In this paper, we extend the vertex nomination problem to a very general statistical model of graphs. Further, drawing inspiration from the long-established classification framework in the pattern recognition literature, we provide definitions for the key notions of Bayes optimality and consistency in our extended vertex nomination framework, including a derivation of the Bayes optimal vertex nomination scheme. In addition, we prove that no universally consistent vertex nomination schemes exist. Illustrative examples are provided throughout.


RACE: Large-scale ReAding Comprehension Dataset From Examinations

arXiv.org Artificial Intelligence

Collected from the English exams for middle and high school Chinese students in the age range between 12 to 18, RACE consists of near 28,000 passages and near 100,000 questions generated by human experts (English instructors), and covers a variety of topics which are carefully designed for evaluating the students' ability in understanding and reasoning. In particular, the proportion of questions that requires reasoning is much larger in RACE than that in other benchmark datasets for reading comprehension, and there is a significant gap between the performance of the state-of-the-art models (43%) and the ceiling human performance (95%). We hope this new dataset can serve as a valuable resource for research and evaluation in machine comprehension.


Using Options and Covariance Testing for Long Horizon Off-Policy Policy Evaluation

arXiv.org Artificial Intelligence

Evaluating a policy by deploying it in the real world can be risky and costly. Off-policy policy evaluation (OPE) algorithms use historical data collected from running a previous policy to evaluate a new policy, which provides a means for evaluating a policy without requiring it to ever be deployed. Importance sampling is a popular OPE method because it is robust to partial observability and works with continuous states and actions. However, the amount of historical data required by importance sampling can scale exponentially with the horizon of the problem: the number of sequential decisions that are made. We propose using policies over temporally extended actions, called options, and show that combining these policies with importance sampling can significantly improve performance for long-horizon problems. In addition, we can take advantage of special cases that arise due to options-based policies to further improve the performance of importance sampling. We further generalize these special cases to a general covariance testing rule that can be used to decide which weights to drop in an IS estimate, and derive a new IS algorithm called Incremental Importance Sampling that can provide significantly more accurate estimates for a broad class of domains.


Teachers Often Ask Kids to Learn in Ways That Exceed Adult-Sized Attention Spans, Study Finds

U.S. News

Observers are trained to observe children through their peripheral vision so that a child is unaware that he or she is being observed. Observers look at every child, one at a time, in a specified order. As soon as a child is showing a clear behavior, whether on or off task, it is noted and the observer moves on to the next child on his list. More than a dozen observations are taken for each child during each classroom session. This gives equal weight to all the children in the class and avoids overemphasizing attention-grabbing behaviors or highly distractable children.


Robot to Speak at Indiana University's Lecturer Series

U.S. News

Sophia, who became a citizen of Saudi Arabia in October, has a face that can show expression and metal hands. Sophia's clear "skull" shows the inner working wires of the artificially intelligent brain, which functions through a Wi-Fi connection pumped with information and a cohesive vocabulary.


Artificial Intelligence in Law Schools: Busting the Silo

#artificialintelligence

As we further consider how to train future lawyers for the Algorithmic Society and develop the quality of thinking, listening, relating, collaborating, and learning that will define smartness in this new age, law schools must reach beyond their storied walls. In law, we must got beyond talking about algorithmic implications to actually help shape algorithmic performance. We need lawyers and programmers to work together to create a sound "machine learning corpus." There's potential for an entirely new subfield to emerge if given the right support. With many law school attached to major research universities, it's a great place to start this cross-pollination and interdisciplinary work.


Kids coding languages: Why is today's Google Doodle so important for computer programming?

The Independent - Tech

Google Doodles are often challenging, fun and help illuminate the ideas and people that have changed the world. But today's homepage celebrates something even more fundamental than normal. The Doodle centres on coding – helping children to learn the languages that power the site and its homepage itself. And it even lets people do some of that themselves, helping teach some coding basics. It is just one of the various projects by Google and its competitors like Apple that is intended to allow children to code.


MIT scientist shares insights

#artificialintelligence

Hyderabad: ICFAI Foundation for Higher Education (IFHE), deemed to be University of ICFAI organised a lecture on'Big Data and Artificial Intelligence'. Dr Kalyan Veeramachaneni, Principal Research Scientist at the Laboratory for Information and Decision System (LIDS), Massachusetts Institute of Technology delivered the lecture. He shared his experiences about Big Data, Machine Learning, and Artificial Intelligence with the students. Dr. Kalyan said Machine Learning and predictive models are traditionally created by human scientists by generating feature metrics and the generated models are deployed to address various business requirements. He stated that automating the process of creating the models, proved a time-saving initiative.


AI-powered language learning promises to fast-track fluency

#artificialintelligence

A linguistics company is using AI to shorten the time it takes to learn a new language. It takes about 200 hours, using traditional methods, to gain basic proficiency in a new language. This AI-powered platform claims it can teach from beginner to fluency in just a few months – through once-daily 20 minute lessons. Learning a new language is hard. Some people seem to pick up new dialects with ease, but for the rest of us it's a trudge through rote memorization.