Goto

Collaborating Authors

 Asia


Why Learning of Large-Scale Neural Networks Behaves Like Convex Optimization

arXiv.org Artificial Intelligence

In this paper, we present some theoretical work to explain why simple gradient descent methods are so successful in solving non-convex optimization problems in learning large-scale neural networks (NN). After introducing a mathematical tool called canonical space, we have proved that the objective functions in learning NNs are convex in the canonical model space. We further elucidate that the gradients between the original NN model space and the canonical space are related by a pointwise linear transformation, which is represented by the so-called disparity matrix. Furthermore, we have proved that gradient descent methods surely converge to a global minimum of zero loss provided that the disparity matrices maintain full rank. If this full-rank condition holds, the learning of NNs behaves in the same way as normal convex optimization. At last, we have shown that the chance to have singular disparity matrices is extremely slim in large NNs. In particular, when over-parameterized NNs are randomly initialized, the gradient decent algorithms converge to a global minimum of zero loss in probability.


Learning Task Knowledge and its Scope of Applicability in Experience-Based Planning Domains

arXiv.org Artificial Intelligence

Experience-based planning domains (EBPDs) have been recently proposed to improve problem solving by learning from experience. EBPDs provide important concepts for long-term learning and planning in robotics. They rely on acquiring and using task knowledge, i.e., activity schemata, for generating concrete solutions to problem instances in a class of tasks. Using Three-Valued Logic Analysis (TVLA), we extend previous work to generate a set of conditions as the scope of applicability for an activity schema. The inferred scope is a bounded representation of a set of problems of potentially unbounded size, in the form of a 3-valued logical structure, which allows an EBPD system to automatically find an applicable activity schema for solving task problems. We demonstrate the utility of our approach in a set of classes of problems in a simulated domain and a class of real world tasks in a fully physically simulated PR2 robot in Gazebo.


Representative Task Self-selection for Flexible Clustered Lifelong Learning

arXiv.org Artificial Intelligence

Consider the lifelong learning paradigm whose objective is to learn a sequence of tasks depending on previous experiences, e.g., knowledge library or deep network weights. However, the knowledge libraries or deep networks for most recent lifelong learning models are with prescribed size, and can degenerate the performance for both learned tasks and coming ones when facing with a new task environment (cluster). To address this challenge, we propose a novel incremental clustered lifelong learning framework with two knowledge libraries: feature learning library and model knowledge library, called Flexible Clustered Lifelong Learning (FCL3). Specifically, the feature learning library modeled by an autoencoder architecture maintains a set of representation common across all the observed tasks, and the model knowledge library can be self-selected by identifying and adding new representative models (clusters). When a new task arrives, our proposed FCL3 model firstly transfers knowledge from these libraries to encode the new task, i.e., effectively and selectively soft-assigning this new task to multiple representative models over feature learning library. Then, 1) the new task with a higher outlier probability will then be judged as a new representative, and used to redefine both feature learning library and representative models over time; or 2) the new task with lower outlier probability will only refine the feature learning library. For model optimization, we cast this lifelong learning problem as an alternating direction minimization problem as a new task comes. Finally, we evaluate the proposed framework by analyzing several multi-task datasets, and the experimental results demonstrate that our FCL3 model can achieve better performance than most lifelong learning frameworks, even batch clustered multi-task learning models.


Ultra-Scalable Spectral Clustering and Ensemble Clustering

arXiv.org Machine Learning

This paper focuses on scalability and robustness of spectral clustering for extremely large-scale datasets with limited resources. Two novel algorithms are proposed, namely, ultra-scalable spectral clustering (U-SPEC) and ultra-scalable ensemble clustering (U-SENC). In U-SPEC, a hybrid representative selection strategy and a fast approximation method for K-nearest representatives are proposed for the construction of a sparse affinity sub-matrix. By interpreting the sparse sub-matrix as a bipartite graph, the transfer cut is then utilized to efficiently partition the graph and obtain the clustering result. In U-SENC, multiple U-SPEC clusterers are further integrated into an ensemble clustering framework to enhance the robustness of U-SPEC while maintaining high efficiency. Based on the ensemble generation via multiple U-SEPC's, a new bipartite graph is constructed between objects and base clusters and then efficiently partitioned to achieve the consensus clustering result. It is noteworthy that both U-SPEC and U-SENC have nearly linear time and space complexity, and are capable of robustly and efficiently partitioning ten-million-level nonlinearly-separable datasets on a PC with 64GB memory. Experiments on various large-scale datasets have demonstrated the scalability and robustness of our algorithms. The MATLAB code and experimental data are available at https://www.researchgate.net/publication/330760669.


Deep Learning Based Motion Planning For Autonomous Vehicle Using Spatiotemporal LSTM Network

arXiv.org Artificial Intelligence

Motion Planning, as a fundamental technology of automatic navigation for the autonomous vehicle, is still an open challenging issue in the real-life traffic situation and is mostly applied by the model-based approaches. However, due to the complexity of the traffic situations and the uncertainty of the edge cases, it is hard to devise a general motion planning system for the autonomous vehicle. In this paper, we proposed a motion planning model based on deep learning (named as spatiotemporal LSTM network), which is able to generate a real-time reflection based on spatiotemporal information extraction. To be specific, the model based on spatiotemporal LSTM network has three main structure. Firstly, the Convolutional Long-short Term Memory (Conv-LSTM) is used to extract hidden features through sequential image data. Then, the 3D Convolutional Neural Network(3D-CNN) is applied to extract the spatiotemporal information from the multi-frame feature information. Finally, the fully connected neural networks are used to construct a control model for autonomous vehicle steering angle. The experiments demonstrated that the proposed method can generate a robust and accurate visual motion planning results for the autonomous vehicle.


Researchers use AI to predict progression of neurodegenerative diseases: Researchers with Ben-Gurion University of the Negev in Israel have created an artificial intelligence platform for tracking and predicting the progression of neurodegenerative diseases.

#artificialintelligence

Researchers with Ben-Gurion University of the Negev in Israel have created an artificial intelligence platform for tracking and predicting the progression of neurodegenerative diseases. The platform, developed by professor Boaz Lerner of the university's department of industrial engineering and management, will first be used for amyotrophic lateral sclerosis, also called Lou Gehrig's disease. ALS is a fatal neurodegenerative disease that causes death of motor neurons that control voluntary muscles. This muscle atrophy leads to progressive weakness and paralysis, difficulty speaking, swallowing and breathing. The researchers then plan to use the platform for Alzheimer's, Parkinson's and other neurodegenerative diseases.


Destroying life-ending asteroids headed for Earth will be tougher than we thought

Daily Mail - Science & tech

Apocalyptic asteroids heading for Earth may be harder to destroy than the Hollywood sci-fi films would have us believe. Scientists studying just how easy it would be to blow up a life-threatening space rock found they are stronger and more resilient than previously imagined. They say the discovery could aid in the creation of asteroid deflection weapons and for designing efficient asteroid mining techniques. Researchers found the fallout from the enormous collision would be split into two different stages. 'We used to believe that the larger the object, the more easily it would break, because bigger objects are more likely to have flaws,' says Charles El Mir, a recent PhD graduate from Johns Hopkins University, who led the study.


The next "Deep Blue" moment: Self-flying drone racing

#artificialintelligence

In 1997, IBM's "Deep Blue" computer defeated grandmaster Gary Kasparov in a match of chess. It was an historic moment, marking the end of an era where humans could defeat machines in complex strategy games. Today, artificial intelligence (AI) bots can defeat humans in not only chess, but nearly every digital game that exists. However, while we're starting to see some progress with AI-proof-of concepts in motorsports, ping pong and even basketball, AI has yet to come close to beating humans in real-life physical sports. Doing so will require a major technical leap from today's state-of-the-art AI technology, advancing it to a place where AI can interact with, and make sense of, the physical world and unknown conditions, including physical contact from fellow racers or players โ€“ all while navigating a game strategy, race course, set of rules and other complex challenges.


China's Huawei has big ambitions to weaken the US grip on AI leadership

MIT Technology Review

Ren Zhengfei, the reclusive founder and CEO of China's embattled tech giant, Huawei, is defiant about American efforts to impede his company with lawsuits and restrictions. "There is no way the US can crush us," Ren said in a rare recent interview with international media. "The world cannot leave us because we are more advanced." It might sound like bluff and bluster, but these words carry a measure of truth. Huawei's technology road map, especially in the field of artificial intelligence, points to a company that is progressing more rapidly--and on more technology fronts--than any other business in the world.


John Oliver Has Not Been Replaced by a Robot (Yet)

Slate

Despite what Donald Trump would have you believe, the biggest factor when it comes to American employment is automation, not job theft by Mexico or China or other foreign countries that the president says "you've never even heard of." Although as John Oliver points out, Trump is the same person who reportedly pronounced Nepal and Bhutan as nipple and button, so the list of countries he's never heard of might be higher than average. Elsewhere in the segment, Oliver stopped listing fake countries long enough to explain in detail how machines are replacing jobs in some fields and how that can actually a good thing (unless you want to kill a lumberjack). He also broke the news to some kids who will probably grow up to do jobs that don't already exist, like "crypto-baker" or "snail rehydrater." Good thing that unlike "mermaid doctor," the job of "culture blogger" will never be replaced by BEEP BOOP ERROR 404.