Deep Learning
Deep Sentiment Analysis using a Graph-based Text Representation
Bijari, Kayvan, Zare, Hadi, Veisi, Hadi, Kebriaei, Emad
Accordingly, a prime step in text mining applications is to extract interesting patterns and features, from this supply of unstructured data. Feature extraction can be considered as the core of social media mining tasks such as sentiment analysis, event detection, and news recommendation [2]. In the literature, sentiment analysis tends to be used to refer to the task of classifying the polarity of a given piece of text at the document, sentence, feature, or aspect level [23]. There are various applications on a variety of domains which utilize sentiment analysis, in this regard one can mention applying the sentiment analysis for political reviews to estimate the general viewpoint of the parties [43], predicting stock market prices based on sentiment analysis by utilizing the different financial news data [5], and making use of the sentiment analysis to recognize the current medical and psychological status for a community [23]. Machine learning algorithms and statistical learning techniques have been rising in a variety of scientific fields [9, 10]. A number of machine learning techniques have been proposed to perform the task of sentiment analysis. As one of the powerful sub-domains of machine learning in recent years, deep learning models are emerging as a persuasive computational tool, they have affected many research areas and can be traced in many applications. With respect to the deep learning, textual deep representation models attempt to discover and present intricate syntactic and semantic representations of texts, automatically from data without any handmade feature engineering.
Rethinking Action Spaces for Reinforcement Learning in End-to-end Dialog Agents with Latent Variable Models
Zhao, Tiancheng, Xie, Kaige, Eskenazi, Maxine
Defining action spaces for conversational agents and optimizing their decision-making process with reinforcement learning is an enduring challenge. Common practice has been to use handcrafted dialog acts, or the output vocabulary, e.g. in neural encoder decoders, as the action spaces. Both have their own limitations. This paper proposes a novel latent action framework that treats the action spaces of an end-to-end dialog agent as latent variables and develops unsupervised methods in order to induce its own action space from the data. Comprehensive experiments are conducted examining both continuous and discrete action types and two different optimization methods based on stochastic variational inference. Results show that the proposed latent actions achieve superior empirical performance improvement over previous word-level policy gradient methods on both DealOrNoDeal and MultiWoz dialogs. Our detailed analysis also provides insights about various latent variable approaches for policy learning and can serve as a foundation for developing better latent actions in future research.
Medical Multimodal Classifiers Under Scarce Data Condition
Aydin, Faik, Zhang, Maggie, Ananda-Rajah, Michelle, Haffari, Gholamreza
Data is one of the essential ingredients to power deep learning research. Small datasets, especially specific to medical institutes, bring challenges to deep learning training stage. This work aims to develop a practical deep multimodal that can classify patients into abnormal and normal categories accurately as well as assist radiologists to detect visual and textual anomalies by locating areas of interest. The detection of the anomalies is achieved through a novel technique which extends the integrated gradients methodology with an unsupervised clustering algorithm. This technique also introduces a tuning parameter which trades off true positive signals to denoise false positive signals in the detection process. To overcome the challenges of the small training dataset which only has 3K frontal X-ray images and medical reports in pairs, we have adopted transfer learning for the multimodal which concatenates the layers of image and text submodels. The image submodel was trained on the vast ChestX-ray14 dataset, while the text submodel transferred a pertained word embedding layer from a hospital-specific corpus. Experimental results show that our multimodal improves the accuracy of the classification by 4% and 7% on average of 50 epochs, compared to the individual text and image model, respectively.
Transfer Learning for Non-Intrusive Load Monitoring
DIncecco, Michele, Squartini, Stefano, Zhong, Mingjun
Non-intrusive load monitoring (NILM) is a technique to recover source appliances from only the recorded mains in a household. NILM is unidentifiable and thus a challenge problem because the inferred power value of an appliance given only the mains could not be unique. To mitigate the unidentifiable problem, various methods incorporating domain knowledge into NILM have been proposed and shown effective experimentally. Recently, among these methods, deep neural networks are shown performing best. Arguably, the recently proposed sequence-to-point (seq2point) learning is promising for NILM. However, the results were only carried out on the same data domain. It is not clear if the method could be generalised or transferred to different domains, e.g., the test data were drawn from a different country comparing to the training data. We address this issue in the paper, and two transfer learning schemes are proposed, i.e., appliance transfer learning (ATL) and cross-domain transfer learning (CTL). For ATL, our results show that the latent features learnt by a `complex' appliance, e.g., washing machine, can be transferred to a `simple' appliance, e.g., kettle. For CTL, our conclusion is that the seq2point learning is transferable. Precisely, when the training and test data are in a similar domain, seq2point learning can be directly applied to the test data without fine tuning; when the training and test data are in different domains, seq2point learning needs fine tuning before applying to the test data. Interestingly, we show that only the fully connected layers need fine tuning for transfer learning.
A Degeneracy Framework for Scalable Graph Autoencoders
Salha, Guillaume, Hennequin, Romain, Tran, Viet Anh, Vazirgiannis, Michalis
In this paper, we present a general framework to scale graph autoencoders (AE) and graph variational autoencoders (VAE). This framework leverages graph degeneracy concepts to train models only from a dense subset of nodes instead of using the entire graph. Together with a simple yet effective propagation mechanism, our approach significantly improves scalability and training speed while preserving performance. We evaluate and discuss our method on several variants of existing graph AE and VAE, providing the first application of these models to large graphs with up to millions of nodes and edges. We achieve empirically competitive results w.r.t. several popular scalable node embedding methods, which emphasizes the relevance of pursuing further research towards more scalable graph AE and VAE.
Deep Learning Approach on Information Diffusion in Heterogeneous Networks
Molaei, Soheila, Zare, Hadi, Veisi, Hadi
There are many real-world knowledge based networked systems with multi-type interacting entities that can be regarded as heterogeneous networks including human connections and biological evolutions. One of the main issues in such networks is to predict information diffusion such as shape, growth and size of social events and evolutions in the future. While there exist a variety of works on this topic mainly using a threshold-based approach, they suffer from the local viewpoint on the network and sensitivity to the threshold parameters. In this paper, information diffusion is considered through a latent representation learning of the heterogeneous networks to encode in a deep learning model. To this end, we propose a novel meta-path representation learning approach, Heterogeneous Deep Diffusion(HDD), to exploit meta-paths as main entities in networks. At first, the functional heterogeneous structures of the network are learned by a continuous latent representation through traversing meta-paths with the aim of global end-to-end viewpoint. Then, the well-known deep learning architectures are employed on our generated features to predict diffusion processes in the network. The proposed approach enables us to apply it on different information diffusion tasks such as topic diffusion and cascade prediction. We demonstrate the proposed approach on benchmark network datasets through the well-known evaluation measures. The experimental results show that our approach outperforms the earlier state-of-the-art methods.
When Is Technology Too Dangerous to Release to the Public?
Last week, the nonprofit research group OpenAI revealed that it had developed a new text-generation model that can write coherent, versatile prose given a certain subject matter prompt. However, the organization said, it would not be releasing the full algorithm due to "safety and security concerns." Instead, OpenAI decided to release a "much smaller" version of the model and withhold the data sets and training codes that were used to develop it. If your knowledge of the model, called GPT-2, came solely on headlines from the resulting news coverage, you might think that OpenAI had built a weapons-grade chatbot. A headline from Metro U.K. read, "Elon Musk-Founded OpenAI Builds Artificial Intelligence So Powerful That It Must Be Kept Locked Up for the Good of Humanity."
AI extracts speech bubbles from comic strips
Case in point: Researchers at Google parent company Alphabet's DeepMind recently revealed in an academic paper that they'd developed a system capable of segmenting CT scans with "near-human performance." Now, scientists at the University of Potsdam in Germany have developed an AI segmentation tool for a slightly more cartoony medium: comics. In a paper published on the preprint server Arxiv.org During tests involving a dataset containing speech bubbles with "wiggly tails" and "curved corners," it achieved an F1 score (a measure of a test's accuracy) of 0.94, which the researchers claim is state-of-the-art. "Speech balloons usually consist of a carrier, [a symbolic device used to hold the text,] and a tail connecting the carrier to its root character from which the text emerges. Both tails and carriers come in a variety of shapes, outlines, and degrees of wiggliness," the researchers explain.
AI Weekly: Experts say OpenAI's controversial model is a potential threat to society and science
Last week, OpenAI released GPT-2, a conversational AI system that quickly became controversial. Without domain-specific data, GPT-2 achieves state-of-the-art performance in seven of eight natural language understanding benchmarks for things like reading comprehension and answering questions. A paper and some code were released when the unsupervised model, trained on 40GB of internet text, went public, but the entirety of the model wasn't released due to concerns by its creators about "malicious applications of the technology," alluding to things such as automated generation of fake news. As a result, the wider community cannot fully verify or replicate the results. Some, including Keras deep learning library founder François Chollet, called the OpenAI GPT-2 release (or lack thereof) an irresponsible, fear mongering PR tactic and publicity stunt.
Top AI, Machine Learning and Analytics Trends for 2019 and beyond
The technology landscape has changed drastically over the past few years, with many new concepts, approaches and tools. Artificial Intelligence is no longer hype, it is here to stay. The convergence of artificial intelligence, machine learning, deep learning, data science, big data, blockchain, robotics, cloud and IoT are transforming the way we live and work. The main use of artificial intelligence applications in business is automating the process of decision making. However, as artificial intelligence becomes more sophisticated, concern around the fear of artificial intelligence systems as black boxes grows.