Education
From Languages to Information: Another Great NLP Course from Stanford - KDnuggets
We recently highlighted one of the most acclaimed courses on using deep learning techniques for natural language processing, Stanford's freely available Natural Language Processing with Deep Learning (CS224n). Stanford has another fantastic NLP course which is also freely available online, and which is also taught by a world renowned NLP researcher, academic, and author. The course in question is From Languages to Information (CS124), and it is taught by Dan Jurafsky. Just as with the previous Stanford NLP course profile, let's be clear about a couple of things; first, this isn't a recent occurrence, and the course materials and videos (see below) have been available online for quite some time (the materials were once collected into a Coursera course as well). Second, and possibly more importantly, there is no option to enroll, as this is not a MOOC; it is simply the freely available materials from this world-class course on the foundations of natural language processing.
Explainable Artificial Intelligence: a Systematic Review
This has led to the development of a plethora of domain-dependent and context-specific methods for dealing with the interpretation of machine learning (ML) models and the formation of explanations for humans. Unfortunately, this trend is far from being over, with an abundance of knowledge in the field which is scattered and needs organisation. The goal of this article is to systematically review research works in the field of XAI and to try to define some boundaries in the field. From several hundreds of research articles focused on the concept of explainability, about 350 have been considered for review by using the following search methodology. In a first phase, Google Scholar was queried to find papers related to "explainable artificial intelligence", "explainable machine learning" and "interpretable machine learning". Subsequently, the bibliographic section of these articles was thoroughly examined to retrieve further relevant scientific studies. The first noticeable thing, as shown in figure 2 (a), is the distribution of the publication dates of selected research articles: sporadic in the 70s and 80s, receiving preliminary attention in the 90s, showing raising interest in 2000 and becoming a recognised body of knowledge after 2010. The first research concerned the development of an explanation-based system and its integration in a computer program designed to help doctors make diagnoses [3]. Some of the more recent papers focus on work devoted to the clustering of methods for explainability, motivating the need for organising the XAI literature [4, 5, 6].
The growth and form of knowledge networks by kinesthetic curiosity
Zhou, Dale, Lydon-Staley, David M., Zurn, Perry, Bassett, Danielle S.
Throughout life, we might seek a calling, companions, skills, entertainment, truth, self-knowledge, beauty, and edification. The practice of curiosity can be viewed as an extended and open-ended search for valuable information with hidden identity and location in a complex space of interconnected information. Despite its importance, curiosity has been challenging to computationally model because the practice of curiosity often flourishes without specific goals, external reward, or immediate feedback. Here, we show how network science, statistical physics, and philosophy can be integrated into an approach that coheres with and expands the psychological taxonomies of specific-diversive and perceptual-epistemic curiosity. Using this interdisciplinary approach, we distill functional modes of curious information seeking as searching movements in information space. The kinesthetic model of curiosity offers a vibrant counterpart to the deliberative predictions of model-based reinforcement learning. In doing so, this model unearths new computational opportunities for identifying what makes curiosity curious.
The Secret History of Women in Coding
As a teenager in Maryland in the 1950s, Mary Allen Wilkes had no plans to become a software pioneer -- she dreamed of being a litigator. One day in junior high in 1950, though, her geography teacher surprised her with a comment: "Mary Allen, when you grow up, you should be a computer programmer!" Wilkes had no idea what a programmer was; she wasn't even sure what a computer was. The first digital computers had been built barely a decade earlier at universities and in government labs. By the time she was graduating from Wellesley College in 1959, she knew her legal ambitions were out of reach. Her mentors all told her the same thing: Don't even bother applying to law school. "They said: 'Don't do it.
5 Takeaways from the AI for Healthcare Virtual Conference Udacity
As 40% of people infected with COVID-19 are asymptomatic, if a patient is imaged for an unrelated health concern and doctors can identify COVID-19, we'll be in a much better position. In addition to identifying COVID-19 by viral detection and antibody response, we can also suspect viral infection indirectly through resting heart rate. Dr. Eric Topol explained in the "AI for Healthcare Keynote" that for a flu-like illness, the resting heart rate marker allows us to predict illness throughout the country from a wearable device like a Fitbit or Apple watch. Dr. Topol states that heart rate rises before a fever is present, so even if someone doesn't get a fever or experience symptoms, we can still detect that their body is fighting a virus. "Resting heart rate, with the analytics of AI for healthcare, can predict where an outbreak is likely to happen and that's a topic that doesn't get enough respect because people just think test, test, test and they don't understand that digital surveillance with AI can be very useful," said Dr. Topol. Pulse oximetry in wearable devices can also help us detect the virus's damage to the lungs. Dr. Topol thinks that the way to get ahead of this virus is simple: equip everyone with a wearable device that has a pulse oximeter and collects resting heart rate and body temperature. "Here we are in the US spending trillions of dollars. What we should be thinking about is: what can we arm each person with, so that we can help protect them?"
Artificial Intelligence: A Complete Introduction
As you already know, AI is one of the leading technologies in the world today, and people are talking about it much more than ever. We now can find AI applications every where: from finances, marketing, healthcare, to autonomous vehicles, security, or robotics. However, the domain of AI still lacks of qualified employees while the number of investments in AI is increasing rapidly. Thus, open a great opporturnity for people having a background in this domain. After several years of researching and working in AI, now I'd like to share my knowledge and my experiences to people who want to learn about AI, because I really hope that my small contribution can help many ones find a fast and easy way in learning AI.
CompGuessWhat?!: A Multi-task Evaluation Framework for Grounded Language Learning
Suglia, Alessandro, Konstas, Ioannis, Vanzo, Andrea, Bastianelli, Emanuele, Elliott, Desmond, Frank, Stella, Lemon, Oliver
Approaches to Grounded Language Learning typically focus on a single task-based final performance measure that may not depend on desirable properties of the learned hidden representations, such as their ability to predict salient attributes or to generalise to unseen situations. To remedy this, we present GROLLA, an evaluation framework for Grounded Language Learning with Attributes with three sub-tasks: 1) Goal-oriented evaluation; 2) Object attribute prediction evaluation; and 3) Zero-shot evaluation. We also propose a new dataset CompGuessWhat?! as an instance of this framework for evaluating the quality of learned neural representations, in particular concerning attribute grounding. To this end, we extend the original GuessWhat?! dataset by including a semantic layer on top of the perceptual one. Specifically, we enrich the VisualGenome scene graphs associated with the GuessWhat?! images with abstract and situated attributes. By using diagnostic classifiers, we show that current models learn representations that are not expressive enough to encode object attributes (average F1 of 44.27). In addition, they do not learn strategies nor representations that are robust enough to perform well when novel scenes or objects are involved in gameplay (zero-shot best accuracy 50.06%).
A quest for a fair schedule: The Young Physicists' Tournament
Cechlárová, Katarína, Cseh, Ágnes, Jankó, Zsuzsanna, Kireš, Marián, Miňo, Lukáš
The Young Physicists Tournament is an established team-oriented scientific competition between high school students from 37 countries on 5 continents. The competition consists of scientific discussions called Fights. Three or four teams participate in each Fight, each of whom presents a problem while rotating the roles of Presenter, Opponent, Reviewer, and Observer among them. The rules of a few countries require that each team announce in advance 3 problems they will present at the national tournament. The task of the organizers is to choose the composition of Fights in such a way that each team presents each of its chosen problems exactly once and within a single Fight no problem is presented more than once. Besides formalizing these feasibility conditions, in this paper we formulate several additional fairness conditions for tournament schedules. We show that the fulfillment of some of them can be ensured by constructing suitable edge colorings in bipartite graphs. To find fair schedules, we propose integer linear programs and test them on real as well as randomly generated data.
The Value-Improvement Path: Towards Better Representations for Reinforcement Learning
Dabney, Will, Barreto, André, Rowland, Mark, Dadashi, Robert, Quan, John, Bellemare, Marc G., Silver, David
In value-based reinforcement learning (RL), unlike in supervised learning, the agent faces not a single, stationary, approximation problem, but a sequence of value prediction problems. Each time the policy improves, the nature of the problem changes, shifting both the distribution of states and their values. In this paper we take a novel perspective, arguing that the value prediction problems faced by an RL agent should not be addressed in isolation, but rather as a single, holistic, prediction problem. An RL algorithm generates a sequence of policies that, at least approximately, improve towards the optimal policy. We explicitly characterize the associated sequence of value functions and call it the value-improvement path. Our main idea is to approximate the value-improvement path holistically, rather than to solely track the value function of the current policy. Specifically, we discuss the impact that this holistic view of RL has on representation learning. We demonstrate that a representation that spans the past value-improvement path will also provide an accurate value approximation for future policy improvements. We use this insight to better understand existing approaches to auxiliary tasks and to propose new ones. To test our hypothesis empirically, we augmented a standard deep RL agent with an auxiliary task of learning the value-improvement path. In a study of Atari 2600 games, the augmented agent achieved approximately double the mean and median performance of the baseline agent.
Light-in-the-loop: using a photonics co-processor for scalable training of neural networks
Launay, Julien, Poli, Iacopo, Müller, Kilian, Carron, Igor, Daudet, Laurent, Krzakala, Florent, Gigan, Sylvain
As neural networks grow larger and more complex and data-hungry, training costs are skyrocketing. Especially when lifelong learning is necessary, such as in recommender systems or self-driving cars, this might soon become unsustainable. In this study, we present the first optical co-processor able to accelerate the training phase of digitally-implemented neural networks. We rely on direct feedback alignment as an alternative to backpropagation, and perform the error projection step optically. Leveraging the optical random projections delivered by our co-processor, we demonstrate its use to train a neural network for handwritten digits recognition.