Education
What Ingredients Make for an Effective Crowdsourcing Protocol for Difficult NLU Data Collection Tasks?
Nangia, Nikita, Sugawara, Saku, Trivedi, Harsh, Warstadt, Alex, Vania, Clara, Bowman, Samuel R.
Crowdsourcing is widely used to create data for common natural language understanding tasks. Despite the importance of these datasets for measuring and refining model understanding of language, there has been little focus on the crowdsourcing methods used for collecting the datasets. In this paper, we compare the efficacy of interventions that have been proposed in prior work as ways of improving data quality. We use multiple-choice question answering as a testbed and run a randomized trial by assigning crowdworkers to write questions under one of four different data collection protocols. We find that asking workers to write explanations for their examples is an ineffective stand-alone strategy for boosting NLU example difficulty. However, we find that training crowdworkers, and then using an iterative process of collecting data, sending feedback, and qualifying workers based on expert judgments is an effective means of collecting challenging data. But using crowdsourced, instead of expert judgments, to qualify workers and send feedback does not prove to be effective. We observe that the data from the iterative protocol with expert assessments is more challenging by several measures. Notably, the human--model gap on the unanimous agreement portion of this data is, on average, twice as large as the gap for the baseline protocol data.
Search Methods for Sufficient, Socially-Aligned Feature Importance Explanations with In-Distribution Counterfactuals
Hase, Peter, Xie, Harry, Bansal, Mohit
Feature importance (FI) estimates are a popular form of explanation, and they are commonly created and evaluated by computing the change in model confidence caused by removing certain input features at test time. For example, in the standard Sufficiency metric, only the top-k most important tokens are kept. In this paper, we study several under-explored dimensions of FI-based explanations, providing conceptual and empirical improvements for this form of explanation. First, we advance a new argument for why it can be problematic to remove features from an input when creating or evaluating explanations: the fact that these counterfactual inputs are out-of-distribution (OOD) to models implies that the resulting explanations are socially misaligned. The crux of the problem is that the model prior and random weight initialization influence the explanations (and explanation metrics) in unintended ways. To resolve this issue, we propose a simple alteration to the model training process, which results in more socially aligned explanations and metrics. Second, we compare among five approaches for removing features from model inputs. We find that some methods produce more OOD counterfactuals than others, and we make recommendations for selecting a feature-replacement function. Finally, we introduce four search-based methods for identifying FI explanations and compare them to strong baselines, including LIME, Integrated Gradients, and random search. On experiments with six diverse text classification datasets, we find that the only method that consistently outperforms random search is a Parallel Local Search that we introduce. Improvements over the second-best method are as large as 5.4 points for Sufficiency and 17 points for Comprehensiveness. All supporting code is publicly available at https://github.com/peterbhase/ExplanationSearch.
Student Performance Prediction Using Dynamic Neural Models
Delianidi, Marina, Diamantaras, Konstantinos, Chrysogonidis, George, Nikiforidis, Vasileios
We address the problem of predicting the correctness of the student's response on the next exam question based on their previous interactions in the course of their learning and evaluation process. We model the student performance as a dynamic problem and compare the two major classes of dynamic neural architectures for its solution, namely the finite-memory Time Delay Neural Networks (TDNN) and the potentially infinite-memory Recurrent Neural Networks (RNN). Since the next response is a function of the knowledge state of the student and this, in turn, is a function of their previous responses and the skills associated with the previous questions, we propose a two-part network architecture. The first part employs a dynamic neural network (either TDNN or RNN) to trace the student knowledge state. The second part applies on top of the dynamic part and it is a multi-layer feed-forward network which completes the classification task of predicting the student response based on our estimate of the student knowledge state. Both input skills and previous responses are encoded using different embeddings. Regarding the skill embeddings we tried two different initialization schemes using (a) random vectors and (b) pretrained vectors matching the textual descriptions of the skills. Our experiments show that the performance of the RNN approach is better compared to the TDNN approach in all datasets that we have used. Also, we show that our RNN architecture outperforms the state-of-the-art models in four out of five datasets. It is worth noting that the TDNN approach also outperforms the state of the art models in four out of five datasets, although it is slightly worse than our proposed RNN approach. Finally, contrary to our expectations, we find that the initialization of skill embeddings using pretrained vectors offers practically no advantage over random initialization.
Graph-based Exercise- and Knowledge-Aware Learning Network for Student Performance Prediction
Liu, Mengfan, Shao, Pengyang, Zhang, Kun
Predicting student performance is a fundamental task in Intelligent Tutoring Systems (ITSs), by which we can learn about students' knowledge level and provide personalized teaching strategies for them. Researchers have made plenty of efforts on this task. They either leverage educational psychology methods to predict students' scores according to the learned knowledge proficiency, or make full use of Collaborative Filtering (CF) models to represent latent factors of students and exercises. However, most of these methods either neglect the exercise-specific characteristics (e.g., exercise materials), or cannot fully explore the high-order interactions between students, exercises, as well as knowledge concepts. To this end, we propose a Graph-based Exercise- and Knowledge-Aware Learning Network for accurate student score prediction. Specifically, we learn students' mastery of exercises and knowledge concepts respectively to model the two-fold effects of exercises and knowledge concepts. Then, to model the high-order interactions, we apply graph convolution techniques in the prediction process. Extensive experiments on two real-world datasets prove the effectiveness of our proposed Graph-EKLN.
Reinforced Iterative Knowledge Distillation for Cross-Lingual Named Entity Recognition
Liang, Shining, Gong, Ming, Pei, Jian, Shou, Linjun, Zuo, Wanli, Zuo, Xianglin, Jiang, Daxin
Named entity recognition (NER) is a fundamental component in many applications, such as Web Search and Voice Assistants. Although deep neural networks greatly improve the performance of NER, due to the requirement of large amounts of training data, deep neural networks can hardly scale out to many languages in an industry setting. To tackle this challenge, cross-lingual NER transfers knowledge from a rich-resource language to languages with low resources through pre-trained multilingual language models. Instead of using training data in target languages, cross-lingual NER has to rely on only training data in source languages, and optionally adds the translated training data derived from source languages. However, the existing cross-lingual NER methods do not make good use of rich unlabeled data in target languages, which is relatively easy to collect in industry applications. To address the opportunities and challenges, in this paper we describe our novel practice in Microsoft to leverage such large amounts of unlabeled data in target languages in real production settings. To effectively extract weak supervision signals from the unlabeled data, we develop a novel approach based on the ideas of semi-supervised learning and reinforcement learning. The empirical study on three benchmark data sets verifies that our approach establishes the new state-of-the-art performance with clear edges. Now, the NER techniques reported in this paper are on their way to become a fundamental component for Web ranking, Entity Pane, Answers Triggering, and Question Answering in the Microsoft Bing search engine. Moreover, our techniques will also serve as part of the Spoken Language Understanding module for a commercial voice assistant. We plan to open source the code of the prototype framework after deployment.
Artificial Intelligence for Trading
Demand for quantitative talent is growing at incredible rates. Data-driven traders are now responsible for more than 30% of all US stock trades by investors (or about $1 trillion USD worth of investments, up from 14% in 2013). This scenario represents incredible opportunity for individuals eager to apply cutting-edge technologies to trading and finance. Whether you want to pursue a new job in finance, launch yourself on the path to a quant trading career, or master the latest AI applications in trading and quantitative finance, this program will give you the opportunity to build an impressive portfolio of real-world projects. You will build financial models on real data, and work on your own trading strategies using natural language processing, recurrent neural networks, and random forests.
This is a great moment to look for a new job in Artificial Intelligence.
Artificial intelligence (AI) is essential because it allows the software to perform human capacities such as understanding, reasoning, planning, communication, and perception in an increasingly effective, efficient, and low-cost manner. In most business sectors, automating these skills opens up new opportunities. With the significant evolution of algorithms, AI is already a reality. Deep Learning algorithms such as Convolutional Neural Networks (CNNs), for example, have significantly improved computers' ability to recognize objects in images. In addition, Recurrent Neural Networks (RNNs) algorithms produce voice recognition systems that outperform humans.
This is a great moment to look for a new job in Artificial Intelligence.
Artificial intelligence (AI) is essential because it allows the software to perform human capacities such as understanding, reasoning, planning, communication, and perception in an increasingly effective, efficient, and low-cost manner. In most business sectors, automating these skills opens up new opportunities. With the significant evolution of algorithms, AI is already a reality. Deep Learning algorithms such as Convolutional Neural Networks (CNNs), for example, have significantly improved computers' ability to recognize objects in images. In addition, Recurrent Neural Networks (RNNs) algorithms produce voice recognition systems that outperform humans.
Eye-tracking software could make video calls feel more lifelike
A system that tracks your eye movements could help make video calls truer to life. Shlomo Dubnov at the University of California, San Diego (UCSD), was frustrated by the inability to smoothly teach an online music class during the coronavirus pandemic. "With the online setting, we miss a lot of these little non-verbal body gestures and communications," he says. With Ross Greer, a colleague at UCSD, he developed a machine learning system that monitors a presenter's eye movements to track who they are …