Goto

Collaborating Authors

 Asia


fast.ai · Making neural nets uncool again

#artificialintelligence

Today we are launching the 2018 edition of Cutting Edge Deep Learning for Coders, part 2 of fast.ai's Just as with our part 1 Practical Deep Learning for Coders, there are no pre-requisites beyond high school math and 1 year of coding experience--we teach you everything else you need along the way. This course contains all new material, including new state of the art results in NLP classification (up to 20% better than previously known approaches), and shows how to replicate recent record-breaking performance results on Imagenet and CIFAR10. The main libraries used are PyTorch and fastai (we explain why we use PyTorch and why we created the fastai library in this article). Each of the seven lessons includes a video that's around two hours long, an interactive Jupyter notebook, and a dedicated discussion thread on the fast.ai


Diffusion Based Network Embedding

arXiv.org Machine Learning

In network embedding, random walks play a fundamental role in preserving network structures. However, random walk based embedding methods have two limitations. First, random walk methods are fragile when the sampling frequency or the number of node sequences changes. Second, in disequilibrium networks such as highly biases networks, random walk methods often perform poorly due to the lack of global network information. In order to solve the limitations, we propose in this paper a network diffusion based embedding method. To solve the first limitation, our method employs a diffusion driven process to capture both depth information and breadth information. The time dimension is also attached to node sequences that can strengthen information preserving. To solve the second limitation, our method uses the network inference technique based on cascades to capture the global network information. To verify the performance, we conduct experiments on node classification tasks using the learned representations. Results show that compared with random walk based methods, diffusion based models are more robust when samplings under each node is rare. We also conduct experiments on a highly imbalanced network. Results shows that the proposed model are more robust under the biased network structure.


Textual Membership Queries

arXiv.org Machine Learning

Human labeling of textual data can be very time-consuming and expensive, yet it is critical for the success of an automatic text classification system. In order to minimize human labeling efforts, we propose a novel active learning (AL) solution, that does not rely on existing sources of unlabeled data. It uses a small amount of labeled data as the core set for the synthesis of useful membership queries (MQs) - unlabeled instances synthesized by an algorithm for human labeling. Our solution uses modification operators, functions from the instance space to the instance space that change the input to some extent. We apply the operators on the core set, thus creating a set of new membership queries. Using this framework, we look at the instance space as a search space and apply search algorithms in order to create desirable MQs. We implement this framework in the textual domain. The implementation includes using methods such as WordNet and Word2vec, for replacing text fragments from a given sentence with semantically related ones. We test our framework on several text classification tasks and show improved classifier performance as more MQs are labeled and incorporated into the training set. To the best of our knowledge, this is the first work on membership queries in the textual domain.


An $O(N)$ Sorting Algorithm: Machine Learning Sorting

arXiv.org Machine Learning

Sorting, as a fundamental operation on data, has attracted intensively interest from the beginning of computing [1]. Lots of excellent algorithms have been designed, however, it's been proven that sorting algorithms based on comparison have a fundamental requirement of Ω(N log N) comparisons, which means O(N log N) time complexity. Recent years, with the emergence of big data (even terabytes of data), efficiency becomes more important for data processing, and researchers have put many efforts to make the sorting algorithms more efficient. Most of the state-of-art sorting algorithms employ parallel computing to handle big datasets and have accomplished outstanding achievements [2-6]. For example [7], in 2015, FuxiSort [8], developed by Alibaba Group, is a distributed sort implementation on top of Apsara. FuxiSort is able to complete the 100TB Daytona GraySort benchmark in 377 seconds on random non-skewed dataset and 510 seconds on skewed dataset, and Indy GraySort benchmark in 329 seconds.


Convex Programming Based Spectral Clustering

arXiv.org Machine Learning

Clustering is a fundamental task in data analysis, and spectral clustering has been recognized as a promising approach to it. Given a graph describing the relationship between data, spectral clustering explores the underlying cluster structure in two stages. The first stage embeds the nodes of the graph into real space, and the second stage groups the embedded nodes into several clusters. The use of the $k$-means method in the grouping stage is currently standard practice. We present a spectral clustering algorithm that uses convex programming in the grouping stage, and study how well it works. The concept behind the algorithm design lies in the following observation. The nodes with the largest degree in each cluster may be found by computing an enclosing ellipsoid for embedded nodes in real space, and the clusters may be identified by using those nodes. We show that the observations are valid, and the algorithm returns clusters to provide the conductance of graph, if the gap assumption, introduced by Peng el al. at COLT 2015, is satisfied. We also give an experimental assessment of the algorithm's performance.


Extracting Action Sequences from Texts Based on Deep Reinforcement Learning

arXiv.org Artificial Intelligence

Extracting action sequences from natural language texts is challenging, as it requires commonsense inferences based on world knowledge. Although there has been work on extracting action scripts, instructions, navigation actions, etc., they require that either the set of candidate actions be provided in advance, or that action descriptions are restricted to a specific form, e.g., description templates. In this paper, we aim to extract action sequences from texts in free natural language, i.e., without any restricted templates, provided the candidate set of actions is unknown. We propose to extract action sequences from texts based on the deep reinforcement learning framework. Specifically, we view "selecting" or "eliminating" words from texts as "actions", and the texts associated with actions as "states". We then build Q-networks to learn the policy of extracting actions and extract plans from the labeled texts. We demonstrate the effectiveness of our approach on several datasets with comparison to state-of-the-art approaches, including online experiments interacting with humans.


Intel Editorial: The U.S. Needs a National Strategy on Artificial Intelligence

#artificialintelligence

WASHINGTON--(BUSINESS WIRE)--The following is an opinion editorial provided by Brian Krzanich, chief executive officer of Intel Corporation. China, India, Japan, France and the European Union are crafting bold plans for artificial intelligence (AI). They see AI as a means to economic growth and social progress. Meanwhile, the U.S. disbanded its AI taskforce in 2016. The U.S. technology sector has long been a driver of global economic growth.


Alexa Is a Bad Dog

Slate

Future Tense is a partnership of Slate, New America, and Arizona State University that examines emerging technologies, public policy, and society. Just who is man's best friend? Who are these creatures that we've welcomed into our homes--who light up when we call their names, fetch things for us (like news updates and groceries), and make us feel better when we're sad (by putting on Beyoncé and allowing us to order food without even moving)? Unfortunately, it turns out, not just to you. There's a new reason to worry about the fact that your smart device might be listening all the time: She might be listening to someone else.


How will Artificial Intelligence and Internet of Things change the world.

#artificialintelligence

Are you finding it difficult to understand trends in Artificial Intelligence (AI) and the Internet of Things (IoT)? Join AI&IoT Summit - 6-7 June 2018, New Delhi http://ai-iotsummit.com/?utm_source y... Know more at- https://geospatialworldforum.org/ai-a... #gis #AI #IoT Video Courtesy- 1. TIA NOW 2. LinkedIn Learning Solutions 3. AT&T Business 4. Qualcomm 5. Prime Minister's Office of Japan 6.


Trump Administration Vows to Maintain U.S. Edge in AI Technology

WSJ.com: WSJD - Technology

WASHINGTON--White House officials promised to keep the U.S. in the lead on emerging artificial-intelligence technologies, despite growing competition from China and worries about potential impacts on American workers. At a White House conference on artificial intelligence, Trump technology adviser Michael Kratsios pledged that the administration would make a priority of advancing artificial-intelligence research, through greater research funding and other steps. "America has been the global leader in AI, and the Trump administration will ensure our great nation remains the global leader in AI," said Mr. Kratsios, deputy assistant to the president for technology policy, according to prepared text of a keynote speech. In addition to increased federal funding, Mr. Kratsios raised the possibility that AI researchers might eventually gain expanded access to the government's network of national labs, and also the government's "vast troves of taxpayer-funded data, in ways that don't compromise privacy or security." Artificial intelligence is software that enables computers to emulate human intelligence, handling tasks such as recognizing and processing images or language in applications including autonomous driving.