Deep Learning
Siamese recurrent networks learn first-order logic reasoning and exhibit zero-shot compositional generalization
Can neural nets learn logic? We approach this classic question with current methods, and demonstrate that recurrent neural networks can learn to recognize first order logical entailment relations between expressions. We define an artificial language in first-order predicate logic, generate a large dataset of sample 'sentences', and use an automatic theorem prover to infer the relation between random pairs of such sentences. We describe a Siamese neural architecture trained to predict the logical relation, and experiment with recurrent and recursive networks. Siamese Recurrent Networks are surprisingly successful at the entailment recognition task, reaching near perfect performance on novel sentences (consisting of known words), and even outperforming recursive networks. We report a series of experiments to test the ability of the models to perform compositional generalization. In particular, we study how they deal with sentences of unseen length, and sentences containing unseen words. We show that set-ups using LSTMs and GRUs obtain high scores on these tests, demonstrating a form of compositionality.
Enhancing Item Response Theory for Cognitive Diagnosis
Cognitive diagnosis is a fundamental and crucial task in many educational applications, e.g., computer adaptive test and cognitive assignments. Item Response Theory (IRT) is a classical cognitive diagnosis method which can provide interpretable parameters (i.e., student latent trait, question discrimination, and difficulty) for analyzing student performance. However, traditional IRT ignores the rich information in question texts, cannot diagnose knowledge concept proficiency, and it is inaccurate to diagnose the parameters for the questions which only appear several times. To this end, in this paper, we propose a general Deep Item Response Theory (DIRT) framework to enhance traditional IRT for cognitive diagnosis by exploiting semantic representation from question texts with deep learning. In DIRT, we first use a proficiency vector to represent students' proficiency in knowledge concepts and embed question texts and knowledge concepts to dense vectors by Word2Vec. Then, we design a deep diagnosis module to diagnose parameters in traditional IRT by deep learning techniques. Finally, with the diagnosed parameters, we input them into the logistic-like formula of IRT to predict student performance. Extensive experimental results on real-world data clearly demonstrate the effectiveness and interpretation power of DIRT framework.
Human-like machine thinking: Language guided imagination
Human thinking requires the brain to understand the meaning of language and to properly organize the thoughts flow using the language. However, current natural language processing models are primarily limited to the word probability estimation. Here, we proposed a Language Guided Imagination (LGI) network to incrementally learn the meaning and usage of diverse words and syntaxes, aiming to form a humanlike machine thinking process. LGI contains three subsystems: (1) vision system that contains an encoder to disentangle the input or imagined scenarios into abstract population representations, and an imagination decoder to reconstruct imagined scenarios from higher level representations; (2) Language system, which consists of a binarizer to transfer symbol texts into binary vectors, an IPS (mimicking the human IntraParietal Sulcus, implemented by an LSTM) to extract the quantity information from the input texts, and a textizer to convert binary vectors into text symbols; (3) a PFC (mimicking the human PreFrontal Cortex, implemented by an LSTM) that combines inputs in the forms of both language and vision, and predict text symbols and manipulated images accordingly. In this work, the proposed LGI network illustrates the ability to incrementally learn eight different syntaxes and form a machine thinking loop that enables interactions between language and vision system, which hasn't been demonstrated before. The paper presents a new architecture that allows the machine to learn, understand and use language in a humanlike way, which might ultimately enable a machine to construct fictitious'mental' scenario and possess intelligence.
DeepMind Can Now Beat Us at Multiplayer Games, Too
"They can adapt to teammates with arbitrary skills," said Wojciech Czarnecki, a researcher with DeepMind, a lab owned by the same parent company as Google. Through thousands of hours of game play, the agents learned very particular skills, like racing toward the opponent's home base when a teammate was on the verge of capturing a flag. As human players know, the moment the opposing flag is brought to one's home base, a new flag appears at the opposing base, ripe for the taking. DeepMind's project is part of a broad effort to build artificial intelligence that can play enormously complex, three-dimensional video games, including Quake III, Dota 2 and StarCraft II. Many researchers believe that success in the virtual arena will eventually lead to automated systems with improved abilities in the real world.
Is AI research headed in the right direction?
This article is part of Demystifying AI, a series of posts that (try to) disambiguate the jargon and myths surrounding artificial intelligence. In 1956, researchers at Dartmouth College coined the term "artificial intelligence," a field of science that aims to enable machines to replicate the capabilities of the human mind. AI pioneers believed at the time that in short time, "machines will be capableโฆ of doing any work a man can do." For decades, AI scientists and researchers have been trying to recreate the logic and functionalities of the human brain. And for decades, they have dismayed themselves and the general public.
Cloud AI Isn't About Outsourcing Compute It Is About Joining The Front Row Of The AI Revolution
The commercial cloud that was once synonymous with the mundane task of hardware outsourcing has increasingly become far more about the services and analytic capabilities that can be built when computing power is no longer a limitation. Nowhere is this more apparent than in the world of deep learning. Companies like Google not only provide access to bleeding edge battle-tested hardware, but wrap those systems with point-and-click and data-scale analytic offerings that are increasingly democratizing access to AI. Deep learning in the cloud today is no longer merely about outsourcing compute, but actually joining the front row of the AI revolution itself. The early days of the commercial cloud were largely relegated to the unglamorous tasks of migrating workloads from bare metal on-premises computer racks to virtualized remote managed data centers. The focus was often on lifting and transferring applications to data centers that provided scalability and reliability nearly unheard of in on-premises environments.
r/MachineLearning - DeepMind's new neural network model beats AlexNet with 13 images per class
Definitely does have "echos" of BERT and friends from the NLP side of things, though still has a while to go to reach a similarly large revolution in performance. However, the OP's title does not match the claims of the paper. With unsupervised pretraining on over 1M images 13 labels per class they get 64% top-5 accuracy, well below Alexnet's 82% accuracy. While the paper's investigation is pretty thorough, I don't think they mention either compute requirements (given it's deepmind, I would default to assuming it's gigantic) or how the approach scales with different amounts of unsupervised data. Like how does it perform if only training the CPC feature extractor on half of imagenet?
r/deeplearning - Free cloud GPU credits for deep learning
I am working on https://www.tensorpad.com/ Part of our computational capacity is idle; hence, we're offering credits at a free and discounted rate, so that data scientists can benefit from the resources available, and work on neural networks. You can access the free credits by signing up (https://dashboard.tensorpad.com/ and redeeming "promo450" promo code in the Billing tab (https://dashboard.tensorpad.com/billing). For any questions, please contact us here, through support@tensorpad.com, or the Intercom on the site.
r/deeplearning - Hands-on Graph Neural Networks with PyTorch & PyTorch Geometric
Recently, Graph Neural Networks have gained increasing attention from the Machine Learning researchers and the community. With its strong expressiveness, they are likely to be the next game-changing Neural Networks. PyTorch Geometric is one of the fastest Graph Neural Networks frameworks in the world. This article about the basic usage of PyTorch Geometric and how to use it on real-world data.
Webinar - AI enabled data extraction for financial documents
San Jose, CA, based Infrrd Inc is a leading Machine Intelligence partner to Banking, Financial Services and Insurance industries across the globe. Infrrd's focus is on providing AI as a service and leveraging its homegrown machine learning platform to solve analytics and automation related problems for the customers. Infrrd's platforms and algorithms extract deep insights from big data based on artificial intelligence and deep learning and offer these insights to drive decisions & automate extraction for customers.