Education
elBERto: Self-supervised Commonsense Learning for Question Answering
Zhan, Xunlin, Li, Yuan, Dong, Xiao, Liang, Xiaodan, Hu, Zhiting, Carin, Lawrence
Commonsense question answering requires reasoning about everyday situations and causes and effects implicit in context. Typically, existing approaches first retrieve external evidence and then perform commonsense reasoning using these evidence. In this paper, we propose a Self-supervised Bidirectional Encoder Representation Learning of Commonsense (elBERto) framework, which is compatible with off-the-shelf QA model architectures. The framework comprises five self-supervised tasks to force the model to fully exploit the additional training signals from contexts containing rich commonsense. The tasks include a novel Contrastive Relation Learning task to encourage the model to distinguish between logically contrastive contexts, a new Jigsaw Puzzle task that requires the model to infer logical chains in long contexts, and three classic SSL tasks to maintain pre-trained models language encoding ability. On the representative WIQA, CosmosQA, and ReClor datasets, elBERto outperforms all other methods, including those utilizing explicit graph reasoning and external knowledge retrieval. Moreover, elBERto achieves substantial improvements on out-of-paragraph and no-effect questions where simple lexical similarity comparison does not help, indicating that it successfully learns commonsense and is able to leverage it when given dynamic context.
PSP: Pre-trained Soft Prompts for Few-Shot Abstractive Summarization
Liu, Xiaochen, Gao, Yang, Bai, Yu, Li, Jiawei, Hu, Yinan, Huang, Heyan, Chen, Boxing
Few-shot abstractive summarization has become a challenging task in natural language generation. To support it, we designed a novel soft prompts architecture coupled with a prompt pre-training plus fine-tuning paradigm that is effective and tunes only extremely light parameters. The soft prompts include continuous input embeddings across an encoder and a decoder to fit the structure of the generation models. Importantly, a novel inner-prompt placed in the text is introduced to capture document-level information. The aim is to devote attention to understanding the document that better prompts the model to generate document-related content. The first step in the summarization procedure is to conduct prompt pre-training with self-supervised pseudo-data. This teaches the model basic summarizing capabilities. The model is then fine-tuned with few-shot examples. Experimental results on the CNN/DailyMail and XSum datasets show that our method, with only 0.1% of the parameters, outperforms full-model tuning where all model parameters are tuned. It also surpasses Prompt Tuning by a large margin and delivers competitive results against Prefix-Tuning with 3% of the parameters.
Modular Approach to Machine Reading Comprehension: Mixture of Task-Aware Experts
Rayasam, Anirudha, Kamath, Anusha, Kalejaiye, Gabriel Bayomi Tinoco
In this work we present a Mixture of Task-Aware Experts Network for Machine Reading Comprehension on a relatively small dataset. We particularly focus on the issue of common-sense learning, enforcing the common ground knowledge by specifically training different expert networks to capture different kinds of relationships between each passage, question and choice triplet. Moreover, we take inspi ration on the recent advancements of multitask and transfer learning by training each network a relevant focused task. By making the mixture-of-networks aware of a specific goal by enforcing a task and a relationship, we achieve state-of-the-art results and reduce over-fitting.
Reincarnating Reinforcement Learning: Reusing Prior Computation to Accelerate Progress
Agarwal, Rishabh, Schwarzer, Max, Castro, Pablo Samuel, Courville, Aaron, Bellemare, Marc G.
Learning tabula rasa, that is without any prior knowledge, is the prevalent workflow in reinforcement learning (RL) research. However, RL systems, when applied to large-scale settings, rarely operate tabula rasa. Such large-scale systems undergo multiple design or algorithmic changes during their development cycle and use ad hoc approaches for incorporating these changes without re-training from scratch, which would have been prohibitively expensive. Additionally, the inefficiency of deep RL typically excludes researchers without access to industrial-scale resources from tackling computationally-demanding problems. To address these issues, we present reincarnating RL as an alternative workflow or class of problem settings, where prior computational work (e.g., learned policies) is reused or transferred between design iterations of an RL agent, or from one RL agent to another. As a step towards enabling reincarnating RL from any agent to any other agent, we focus on the specific setting of efficiently transferring an existing sub-optimal policy to a standalone value-based RL agent. We find that existing approaches fail in this setting and propose a simple algorithm to address their limitations. Equipped with this algorithm, we demonstrate reincarnating RL's gains over tabula rasa RL on Atari 2600 games, a challenging locomotion task, and the real-world problem of navigating stratospheric balloons. Overall, this work argues for an alternative approach to RL research, which we believe could significantly improve real-world RL adoption and help democratize it further. Open-sourced code and trained agents at https://agarwl.github.io/reincarnating_rl.
Sentiment Analysis on Demonetization in India using Apache Spark - Projects Based Learning
In this article, We have explored the Sentiments of People in India during Demonetization. Even by using small data, I could still gain a lot of valuable insights. I have used Spark SQL and Inbuild graphs provided by Databricks. India is the second-most populous country in the world, with over 1.271 billion people, more than a sixth of the world's population. Let us find out the views of different people on the demonetization by analyzing the tweets from Twitter.
Wiggling toward bio-inspired machine intelligence
Juncal Arbelaiz Mugica is a native of Spain, where octopus is a common menu item. However, Arbelaiz appreciates octopus and similar creatures in a different way, with her research into soft-robotics theory. More than half of an octopus' nerves are distributed through its eight arms, each of which has some degree of autonomy. This distributed sensing and information processing system intrigued Arbelaiz, who is researching how to design decentralized intelligence for human-made systems with embedded sensing and computation. At MIT, Arbelaiz is an applied math student who is working on the fundamentals of optimal distributed control and estimation in the final weeks before completing her PhD this fall.
Run from robots at RoboBoston's Robot Block Party
"These are definitely the droids you're looking for," insists MassRobotics about its Robot Block Party this weekend, in that way that people who make robots like to pretend they won't eventually rise up and subjugate all of humanity. In reality, it's pretty cool that kids and adults alike will get to experience robotics up close and personal when more than 30 different companies and universities take to the Seaport with drones, autonomous vehicles, collaborative robots, humanoids, robot dogs, flying bionic birds, and all sorts of other mechanical beings available for first-hand viewing. Part of RoboBoston, a two-day affair that also features a STEM Field Trip Day and a Robotics and AI Technical Career Fair -- both on Friday -- the Robot Block Party will mark the fifth time the robots have taken over. Er, the Seaport, that is. "The world looks to our cluster for innovations and advances in robotics. RoboBoston is a chance to celebrate, recognize and share our robust community," said Tom Ryden, executive director of MassRobotics.
Inside Abu Dhabi's Thriving Artificial Intelligence Scene
As the United Arab Emirates (UAE) continues to transition from an oil-based economy to a more knowledge-based economy, artificial intelligence is expected to become one of the country's key sectors. The World Economic Forum estimates that artificial intelligence will add nearly $16 trillion to the global economy by 2030. The country's capital city boasts a burgeoning startup community, advanced machine-learning research facilities, and world-class educational institutions like the Khalifa University of Science and Technology (KU) and the Mohamed bin Zayed University of Artificial Intelligence (MBZUAI). "These institutions help anchor Abu Dhabi's world-class AI ecosystem," said Dr. Ernesto Damiani, professor and senior director of Khalifa University's Robotics and Intelligent Systems Institute. "The ecosystem is also composed of several large companies that are invested in developing AI applications across a variety of verticals, from healthcare to energy."
La veille de la cybersรฉcuritรฉ
There are no standard practices for building and managing machine learning (ML) applications. As a result, machine learning projects are not well organized, lack reproducibility, and are prone to complete failure in the long run. We need a model that helps us maintain quality, sustainability, robustness, and cost management throughout the ML life cycle. The Cross-Industry Standard Process for the development of Machine Learning applications with Quality assurance methodology (CRISP-ML(Q)) is an upgraded version of CRISP-DM to ensure quality ML products. These phases require constant iteration and exploration for building better solutions.
Nonstationary data stream classification with online active learning and siamese neural networks
Malialis, Kleanthis, Panayiotou, Christos G., Polycarpou, Marios M.
We have witnessed in recent years an ever-growing volume of information becoming available in a streaming manner in various application areas. As a result, there is an emerging need for online learning methods that train predictive models on-the-fly. A series of open challenges, however, hinder their deployment in practice. These are, learning as data arrive in real-time one-by-one, learning from data with limited ground truth information, learning from nonstationary data, and learning from severely imbalanced data, while occupying a limited amount of memory for data storage. We propose the ActiSiamese algorithm, which addresses these challenges by combining online active learning, siamese networks, and a multi-queue memory. It develops a new density-based active learning strategy which considers similarity in the latent (rather than the input) space. We conduct an extensive study that compares the role of different active learning budgets and strategies, the performance with/without memory, the performance with/without ensembling, in both synthetic and real-world datasets, under different data nonstationarity characteristics and class imbalance levels. ActiSiamese outperforms baseline and state-of-the-art algorithms, and is effective under severe imbalance, even only when a fraction of the arriving instances' labels is available. We publicly release our code to the community.