Media
mrdbourke/machine-learning-roadmap
A roadmap connecting many of the most important concepts in machine learning, how to learn them and what tools to use to perform them. See the full interactive version. Many of the materials in this roadmap were inspired by Daniel Formoso's machine learning mindmaps,so if you enjoyed this one, go and check out his. He also has a mindmap specifically for deep learning too.
Question and Answer Test-Train Overlap in Open-Domain Question Answering Datasets
Lewis, Patrick, Stenetorp, Pontus, Riedel, Sebastian
Ideally Open-Domain Question Answering models should exhibit a number of competencies, ranging from simply memorizing questions seen at training time, to answering novel question formulations with answers seen during training, to generalizing to completely novel questions with novel answers. However, single aggregated test set scores do not show the full picture of what capabilities models truly have. In this work, we perform a detailed study of the test sets of three popular open-domain benchmark datasets with respect to these competencies. We find that 60-70% of test-time answers are also present somewhere in the training sets. We also find that 30% of test-set questions have a near-duplicate paraphrase in their corresponding training sets. Using these findings, we evaluate a variety of popular open-domain models to obtain greater insight into what extent they can actually generalize, and what drives their overall performance. We find that all models perform dramatically worse on questions that cannot be memorized from training sets, with a mean absolute performance difference of 63% between repeated and non-repeated data. Finally we show that simple nearest-neighbor models out-perform a BART closed-book QA model, further highlighting the role that training set memorization plays in these benchmarks
aschern at SemEval-2020 Task 11: It Takes Three to Tango: RoBERTa, CRF, and Transfer Learning
Chernyavskiy, Anton, Ilvovsky, Dmitry, Nakov, Preslav
We describe our system for SemEval-2020 Task 11 on Detection of Propaganda Techniques in News Articles. We developed ensemble models using RoBERTa-based neural architectures, additional CRF layers, transfer learning between the two subtasks, and advanced post-processing to handle the multi-label nature of the task, the consistency between nested spans, repetitions, and labels from similar spans in training. We achieved sizable improvements over baseline fine-tuned RoBERTa models, and the official evaluation ranked our system 3rd (almost tied with the 2nd) out of 36 teams on the span identification subtask with an F1 score of 0.491, and 2nd (almost tied with the 1st) out of 31 teams on the technique classification subtask with an F1 score of 0.62.
Assisted Perception: Optimizing Observations to Communicate State
Reddy, Siddharth, Levine, Sergey, Dragan, Anca D.
We aim to help users estimate the state of the world in tasks like robotic teleoperation and navigation with visual impairments, where users may have systematic biases that lead to suboptimal behavior: they might struggle to process observations from multiple sensors simultaneously, receive delayed observations, or overestimate distances to obstacles. While we cannot directly change the user's internal beliefs or their internal state estimation process, our insight is that we can still assist them by modifying the user's observations. Instead of showing the user their true observations, we synthesize new observations that lead to more accurate internal state estimates when processed by the user. We refer to this method as assistive state estimation (ASE): an automated assistant uses the true observations to infer the state of the world, then generates a modified observation for the user to consume (e.g., through an augmented reality interface), and optimizes the modification to induce the user's new beliefs to match the assistant's current beliefs. We evaluate ASE in a user study with 12 participants who each perform four tasks: two tasks with known user biases -- bandwidth-limited image classification and a driving video game with observation delay -- and two with unknown biases that our method has to learn -- guided 2D navigation and a lunar lander teleoperation video game. A different assistance strategy emerges in each domain, such as quickly revealing informative pixels to speed up image classification, using a dynamics model to undo observation delay in driving, identifying nearby landmarks for navigation, and exaggerating a visual indicator of tilt in the lander game. The results show that ASE substantially improves the task performance of users with bandwidth constraints, observation delay, and other unknown biases.
Gibbs Sampling with People
Harrison, Peter M. C., Marjieh, Raja, Adolfi, Federico, van Rijn, Pol, Anglada-Tort, Manuel, Tchernichovski, Ofer, Larrouy-Maestri, Pauline, Jacoby, Nori
A core problem in cognitive science and machine learning is to understand how humans derive semantic representations from perceptual objects, such as color from an apple, pleasantness from a musical chord, or trustworthiness from a face. Markov Chain Monte Carlo with People (MCMCP) is a prominent method for studying such representations, in which participants are presented with binary choice trials constructed such that the decisions follow a Markov Chain Monte Carlo acceptance rule. However, MCMCP's binary choice paradigm generates relatively little information per trial, and its local proposal function makes it slow to explore the parameter space and find the modes of the distribution. Here we therefore generalize MCMCP to a continuous-sampling paradigm, where in each iteration the participant uses a slider to continuously manipulate a single stimulus dimension to optimize a given criterion such as 'pleasantness'. We formulate both methods from a utility-theory perspective, and show that the new method can be interpreted as 'Gibbs Sampling with People' (GSP). Further, we introduce an aggregation parameter to the transition step, and show that this parameter can be manipulated to flexibly shift between Gibbs sampling and deterministic optimization. In an initial study, we show GSP clearly outperforming MCMCP; we then show that GSP provides novel and interpretable results in three other domains, namely musical chords, vocal emotions, and faces. We validate these results through large-scale perceptual rating experiments. The final experiments combine GSP with a state-of-the-art image synthesis network (StyleGAN) and a recent network interpretability technique (GANSpace), enabling GSP to efficiently explore high-dimensional perceptual spaces, and demonstrating how GSP can be a powerful tool for jointly characterizing semantic representations in humans and machines.
GSTS awarded contribution for Space-Based Artificial Intelligence
HALIFAX, NS, Aug. 4, 2020 /CNW/ - Global Spatial Technology Solutions ("GSTS" or "the Company") an Artificial Intelligence (AI) and Maritime Analytics company today announced that it has been selected by the Canadian Space Agency (CSA) to develop space-based AI capability to support enhanced decision-making for a range of space applications focused on tasks using computer vision (such as would be used by exploration landers, rovers, robotics or Earth observation systems). This project is funded under the Space Technology Development Program. "This contribution will enable GSTS to expand our growing AI capabilities into the space sector to support decision making based on the same techniques we utilize in the maritime domain, enabling detection, recognition and prediction," said Richard Kolacz, GSTS CEO. "It is equivalent to placing the brain next to the eyes of any space asset or sensor in order to support decision-making locally, rather than having to relay all the data to Earth for analysis before a decision can be made. It is the first step in the development of truly autonomous space capability." Computer vision involves the automatic extraction, analysis and understanding of information gleaned from digital images. By applying machine learning, which is a type of AI, it can enhance and optimize the production of actionable insights much faster and more accurately than a human can.
Multi-Label Image Classification in TensorFlow 2.0
The 2.2M parameters in MobileNet are frozen, but there are 1.3K trainable parameters in the dense layers. You need to apply the sigmoid activation function in the final neurons to ouput a probability score for each genre apart. By doing so, you are relying on multiple logistic regressions to train simultaneously inside the same model. Every final neuron will act as a seperate binary classifier for one single class, even though the features extracted are common to all final neurons. When generating predictions with this model, you should expect an independant probability score for each genre and that all probability scores do not necessarily sum up to 1. This is different from using a softmax layer in multi-class classification where the sum of probability scores in the output is equal to 1.
TikTok bans deepfakes to fight misinformation and election meddling
TikTok has become the latest platform to ban deepfakes. Under its new policy, the app says it "prohibits synthetic or manipulated content that misleads users by distorting the truth of events in a way that could cause harm." "Our intent is to protect users from things like shallow or deep fakes, so while this kind of content was broadly covered by our guidelines already, this update makes the policy clearer for our users." TikTok's General Manager Vanessa Pappas wrote in a statement. The deepfake ban is part of a broader set of policy changes meant to fight misinformation and election meddling.