Education
An Oral Exam for Measuring a Dialog System’s Capabilities
Cohen, David (Carnegie Mellon University) | Lane, Ian (Carnegie Mellon University)
This paper suggests a model and methodology for measuring the breadth and flexibility of a dialog system's capabilities. The approach relies on having human evaluators administer a targeted oral exam to a system and provide their subjective views of that system's performance on each test problem. We present results from one instantiation of this test being performed on two publicly-accessible dialog systems and a human, and show that the suggested metrics do provide useful insights into the relative strengths and weaknesses of these systems. Results suggest that this approach can be performed with reasonable reliability and with reasonable amounts of effort. We hope that authors will augment their reporting with this approach to improve clarity and make more direct progress toward broadly-capable dialog systems.
Strategyproof Peer Selection: Mechanisms, Analyses, and Experiments
Aziz, Haris (Data61 and University of New South Wales) | Lev, Omer (University of Toronto) | Mattei, Nicholas (Data61 and University of New South Wales) | Rosenschein, Jeffrey S. (The Hebrew University of Jerusalem) | Walsh, Toby (Data61 and University of New South Wales)
We study an important crowdsourcing setting where agents evaluate one another and, based on these evaluations, a subset of agents are selected. This setting is ubiquitous when peer review is used for distributing awards in a team, allocating funding to scientists, and selecting publications for conferences. The fundamental challenge when applying crowdsourcing in these settings is that agents may misreport their reviews of others to increase their chances of being selected. We propose a new strategyproof (impartial) mechanism called Dollar Partition that satisfies desirable axiomatic properties. We then show, using a detailed experiment with parameter values derived from target real world domains, that our mechanism performs better on average, and in the worst case, than other strategyproof mechanisms in the literature.
Scaling-up Empirical Risk Minimization: Optimization of Incomplete U-statistics
Clémençon, Stéphan, Bellet, Aurélien, Colin, Igor
In a wide range of statistical learning problems such as ranking, clustering or metric learning among others, the risk is accurately estimated by $U$-statistics of degree $d\geq 1$, i.e. functionals of the training data with low variance that take the form of averages over $k$-tuples. From a computational perspective, the calculation of such statistics is highly expensive even for a moderate sample size $n$, as it requires averaging $O(n^d)$ terms. This makes learning procedures relying on the optimization of such data functionals hardly feasible in practice. It is the major goal of this paper to show that, strikingly, such empirical risks can be replaced by drastically computationally simpler Monte-Carlo estimates based on $O(n)$ terms only, usually referred to as incomplete $U$-statistics, without damaging the $O_{\mathbb{P}}(1/\sqrt{n})$ learning rate of Empirical Risk Minimization (ERM) procedures. For this purpose, we establish uniform deviation results describing the error made when approximating a $U$-process by its incomplete version under appropriate complexity assumptions. Extensions to model selection, fast rate situations and various sampling techniques are also considered, as well as an application to stochastic gradient descent for ERM. Finally, numerical examples are displayed in order to provide strong empirical evidence that the approach we promote largely surpasses more naive subsampling techniques.
This 24-year-old venture capitalist is using UC Berkeley as his own incubator
Universities have in recent years awakened to the fact that students can help them make money through more than just tuition and board. Stanford University started investing in students' start-ups in 2013. Harvard University does the same through its Xfund. Last year, the University of California launched a 250-million venture fund to invest in companies that grow out of the UC system. Now, UC Berkeley is getting in on the game -- through a new fund led by 24-year-old Los Angeles native Jeremy Fiance.
Securing safe water through Cortana Intelligence Suite
Jacob Katuva used to get up at dawn to cycle 12 miles from his village to collect water with his uncles and cousins when he was growing up in Kenya. Now he is part of a research team at the University of Oxford using cloud computing and mobile sensors to monitor water wells and help ensure that thousands of villages in rural Africa and Asia have a safe, secure supply of water. The time spent finding and carrying water, if local wells are not reliable, steals precious time from farming, making a living or going to school. It can even force people to revert to unsanitary water sources shared with animals. Water issues are tied to a cycle of poverty.
The Data Structures and Algorithms Learning Problem - DZone Big Data
There was more about Foundations of Multidimensional and Metric Data Structures by Hanan Samet being too detailed, Stack Overflow being too high-level, and more hand-wringing after that, too. The email was pleading for some book or series of blog posts that would somehow educate data science folks on more fundamental issues of data structures and algorithms. Perhaps getting them to drop some dimensions when doing k-NN problems or perhaps exploit some other data structure that didn't involve 100's of columns. I'm guessing because -- like a lot of hand-waving emails -- it didn't involve code. If there is a lack of awareness of appropriate data structures, the real place to start is The Algorithm Design Manual by Steven Skiena.
Data, not algorithms, is key to machine learning success
There has been an explosion in machine learning activity, and Shivon Zilis recently mapped out the current machine intelligence ecosystem as we enter 2016. This is one of the key areas that we'll be following this year. While the opportunities here are tremendous, the exuberance surrounding machine learning distracts startups from a key hurdle: it's data, not algorithms, that will dictate who wins in this space. Algorithms have largely been commoditized by now, so a machine learning company built around publicly accessible data isn't defensible. But, startups face a serious chicken and egg problem: they have to convince people to give them data, but the machine intelligence service won't be useful until people (and a lot of people) are actually using the service and sharing their data.
Singularity University: meet the people who are building our future
It's day one at the Singularity University: the opening address has just been delivered by a hologram. Craig Venter, who was one of the first scientists to sequence the human genome and created the first synthetic life form, is up next. And later, we will see two people, paralysed from the waist down, use robotic exoskeletons to rise up and walk. But first, the co-founder of the Singularity University, Peter Diamandis, gives us our instructions for the day. Your task, he says, is to pick one of the "grand challenges of humanity" – the lack of clean drinking water, say. And then come up with an idea that "can positively impact the lives of a billion people". Some of us haven't even had coffee yet. There's about 50 of us present and the room has been divided up into tables, one for education, another for poverty, another for water, and I'm not sure where I should sit. Diane Murphy, the university's PR executive, hesitates for a moment and then directs me over to the table marked "food". "Tell you what," she says.
Root Is a Little Robot on a Mission to Teach Kids to Code
Computing jobs are growing at twice the national rate of other types of employment. By 2020, the Bureau of Labor Statistics says, the US will have 1 million more computer science-related jobs than graduates qualified to fill them. In December, President Obama announced the Computer Science for All Initiative, pledging 4 billion in funding for computer science education in the nation's schools. Yet all kinds of dysfunction keeps the country from closing the deficit in computer science talent, according to a survey by Google and Gallup. Yes, school budgets are a problem, and teachers have a limited time to devote to additional classes.
Be kind to artificial intelligence
Mike Finley is a co-founder of AnswerRocket in charge of natural language processing and machine learning. Big innovations come in unexpected bursts. We grow accustomed to life and work as we know it, until something apparently simple brings about bold change. For example, we used phones for 100 years, but making them mobile transformed the world; we had the Internet for decades before the Web browser put digital education, entertainment and shopping in the hands of billions; and we documented our lives with physical pictures, paper records, CD-ROMs and thumb drives until Jeff Bezos brought us "the cloud." When individual creativity is enhanced by technical ingenuity, new behaviors and capabilities emerge.