Goto

Collaborating Authors

 Asia


Mixed-Order Spectral Clustering for Networks

arXiv.org Machine Learning

Clustering is fundamental for gaining insights from complex networks, and spectral clustering (SC) is a popular approach. Conventional SC focuses on second-order structures (e.g., edges connecting two nodes) without direct consideration of higher-order structures (e.g., triangles and cliques). This has motivated SC extensions that directly consider higher-order structures. However, both approaches are limited to considering a single order. This paper proposes a new Mixed-Order Spectral Clustering (MOSC) approach to model both second-order and third-order structures simultaneously, with two MOSC methods developed based on Graph Laplacian (GL) and Random Walks (RW). MOSC-GL combines edge and triangle adjacency matrices, with theoretical performance guarantee. MOSC-RW combines first-order and second-order random walks for a probabilistic interpretation. We automatically determine the mixing parameter based on cut criteria or triangle density, and construct new structure-aware error metrics for performance evaluation. Experiments on real-world networks show 1) the superior performance of two MOSC methods over existing SC methods, 2) the effectiveness of the mixing parameter determination strategy, and 3) insights offered by the structure-aware error metrics.


PPD: Permutation Phase Defense Against Adversarial Examples in Deep Learning

arXiv.org Machine Learning

Deep neural networks have demonstrated cutting edge performance on various tasks including classification. However, it is well known that adversarially designed imperceptible perturbation of the input can mislead advanced classifiers. In this paper, Permutation Phase Defense (PPD), is proposed as a novel method to resist adversarial attacks. PPD combines random permutation of the image with phase component of its Fourier transform. The basic idea behind this approach is to turn adversarial defense problems analogously into symmetric cryptography, which relies solely on safekeeping of the keys for security. In PPD, safe keeping of the selected permutation ensures effectiveness against adversarial attacks. Testing PPD on MNIST and CIFAR-10 datasets yielded state-of-the-art robustness against the most powerful adversarial attacks currently available.


Parallel Clustering of Single Cell Transcriptomic Data with Split-Merge Sampling on Dirichlet Process Mixtures

arXiv.org Machine Learning

Motivation: With the development of droplet based systems, massive single cell transcriptome data has become available, which enables analysis of cellular and molecular processes at single cell resolution and is instrumental to understanding many biological processes. While state-of-the-art clustering methods have been applied to the data, they face challenges in the following aspects: (1) the clustering quality still needs to be improved; (2) most models need prior knowledge on number of clusters, which is not always available; (3) there is a demand for faster computational speed. Results: We propose to tackle these challenges with Parallel Split Merge Sampling on Dirichlet Process Mixture Model (the Para-DPMM model). Unlike classic DPMM methods that perform sampling on each single data point, the split merge mechanism samples on the cluster level, which significantly improves convergence and optimality of the result. The model is highly parallelized and can utilize the computing power of high performance computing (HPC) clusters, enabling massive clustering on huge datasets. Experiment results show the model outperforms current widely used models in both clustering quality and computational speed. Availability: Source code is publicly available on https://github.com/tiehangd/Para_DPMM/tree/master/Para_DPMM_package


Science vs. the state: a family saga at the Caltech of China

MIT Technology Review

On a hot late-summer day in 2005, I sat in a packed, agreeably air-conditioned auditorium and listened to a university administrator welcome the class of 2009. As the popular saying goes, 'The rich go to Peking U, the poor go to Tsinghua, and the ones willing to work themselves to death come to USTC.'" If Peking University is China's Harvard, and Tsinghua is China's MIT, the University of Science and Technology of China, or USTC, is known as "the Caltech of China" for its small size and intense focus on science and engineering. I was proud to be there. But my pride shifted to awkwardness after the speech, when we stood to sing the university anthem, which ends with an exhortation: "Always learn from the people, and learn from the great leader Mao Zedong!" Hearing Mao's name left a bitter taste. It reminded me of career paths my country had denied me. Without the rule of law, I could not become a lawyer. Without a free press, I could not become a journalist. Without democratic elections, I could not become a politician. Instead I did what was expected of Chinese students without political connections or financial resources but with impeccable grades: I came to USTC to study science. The lyrics of the anthem brought up a question my classmates and I would often ponder: Must scientific research be in service of one's country--or can the pursuit of knowledge transcend nationalism? Generations of scientists at USTC have sought to answer this question. The university gave birth to both China's first satellite, launched in 1970, and the world's first quantum-communication satellite, launched in 2016. It is home to China's first synchrotron particle accelerator, and it will soon host a new multibillion-dollar quantum-science center. Over the years, faculty and students have, at times, wielded the university's scientific prestige as a shield to protect academic freedom and political independence. But if the university's rising trajectory in recent years is any indication, science in China thrives most when it serves the state. Today I live and work in the United States. I spoke to many old schoolmates and current USTC researchers to report this article. The story of USTC that emerges reveals the limits of science's ability to transcend China's authoritarian politics. It is also the story of my family across three generations. USTC was founded in Beijing in 1958, to train scientists for China's fledgling nuclear and space programs. Members of the faculty were drawn from China's scientific elite. Fang Lizhi, one of the first, came to teach physics after being deemed too politically outspoken to work on the bomb. "He was actually happy about it!


More than 4,000 jobs in Artificial intelligence lying vacant: Study

#artificialintelligence

A study on the Indian artificial intelligence (AI) industry by Great Learning, the online education company, indicates there are over 4,000 positions related to AI in India that remain vacant due to shortage of qualified talent at mid and senior levels. This is despite the industry growing by 30% in the last one year to $230 million. These opportunities do not include the slew of new jobs advertised every month, but refer to opportunities that have been vacant for a period of 12 months. The 10 leading companies with the most number of AI openings in India this year are: IBM, Accenture, Amazon, Fractal Analytics, Societe Generale, SAP Labs, 24/7 Customer, Atos, Nvidia, Tech Mahindra. When it comes to remuneration, the median salary of AI professionals in India is Rs 14.3 lakh across all experience level and skill-sets, 40% of AI professionals have an entry-level salary of Rs 6 lakh onwards, and 4% command a salary higher than Rs 50 lakh.


ValueLabs Moves Beyond Employee Engagement - HR Transforms to Employee Success Organisation

#artificialintelligence

ValueLabs is upping the employee engagement game and transforming its HR into the'Employee Success Organisation'. All business partners from HR will now wear a new avatar – that of an'Employee Success Partner'. As the name suggests, it is all about enabling the success of every single employee in the company. The company has taken multiple steps to listen to the voice of the employees and fulfil their aspirations. The CEO of the company, Arjun Rao, conducted a'Listening Tour' over several weeks, when he had heart to heart conversations with all the employees in the company.


2018 in science: Climate change, space exploration and water bears

The Japan Times

In casting an eye back over memorable science and environment stories from Japan in 2018, it is impossible to ignore the extreme weather that hit the country. In late June to July, unusually heavy rainfall caused extensive flooding in southwest parts of the country. Some 8 million people were advised to evacuate, more than 200 people died and the floods caused more than ¥1 trillion in damages. Then, later in July, Japan was baked in a debilitating heatwave. The town of Kumagaya, Saitama Prefecture, racked up the highest temperature on record, reaching 41.1 degrees. Across the country, by the end of the heatwave some 125 people had died and more than 57,000 were taken to hospital.


VMAV-C: A Deep Attention-based Reinforcement Learning Algorithm for Model-based Control

arXiv.org Artificial Intelligence

Recent breakthroughs in Go play and strategic games have witnessed the great potential of reinforcement learning in intelligently scheduling in uncertain environment, but some bottlenecks are also encountered when we generalize this paradigm to universal complex tasks. Among them, the low efficiency of data utilization in model-free reinforcement algorithms is of great concern. In contrast, the model-based reinforcement learning algorithms can reveal underlying dynamics in learning environments and seldom suffer the data utilization problem. To address the problem, a model-based reinforcement learning algorithm with attention mechanism embedded is proposed as an extension of World Models in this paper. We learn the environment model through Mixture Density Network Recurrent Network(MDN-RNN) for agents to interact, with combinations of variational auto-encoder(VAE) and attention incorporated in state value estimates during the process of learning policy. In this way, agent can learn optimal policies through less interactions with actual environment, and final experiments demonstrate the effectiveness of our model in control problem.


Deep Representation Learning for Clustering of Health Tweets

arXiv.org Machine Learning

Twitter has been a prominent social media platform for mining population-level health data and accurate clustering of health-related tweets into topics is important for extracting relevant health insights. In this work, we propose deep convolutional autoencoders for learning compact representations of health-related tweets, further to be employed in clustering. We compare our method to several conventional tweet representation methods including bag-of-words, term frequency-inverse document frequency, Latent Dirichlet Allocation and Non-negative Matrix Factorization with 3 different clustering algorithms. Our results show that the clustering performance using proposed representation learning scheme significantly outperforms that of conventional methods for all experiments of different number of clusters. In addition, we propose a constraint on the learned representations during the neural network training in order to further enhance the clustering performance. All in all, this study introduces utilization of deep neural network-based architectures, i.e., deep convolutional autoencoders, for learning informative representations of health-related tweets.


Multiple Sclerosis Lesion Inpainting Using Non-Local Partial Convolutions

arXiv.org Machine Learning

Multiple sclerosis (MS) is an inflammatory demyelinating disease of the central nervous system (CNS) that results in focal injury to the grey and white matter. The presence of white matter lesions biases morphometric analyses such as registration, individual longitudinal measurements and tissue segmentation for brain volume measurements. Lesion-inpainting with intensities derived from surround healthy tissue represent one approach to alleviate such problems. However, existing methods inpaint lesions based on texture information derived from local surrounding tissue, often leading to inconsistent inpainting and the generation of artifacts such as intensity discrepancy and blurriness. Based on these observations, we propose non-local partial convolutions (NLPC) which integrates a Unet-like network with the non-local module. The non-local module is exploited to capture long range dependencies between the lesion area and remaining normal-appearing brain regions. Then, the lesion area is filled by referring to normal-appearing regions with more similar features. This method generates inpainted regions that appear more realistic and natural. Our quantitative experimental results also demonstrate superiority of this technique of existing state-of-the-art inpainting methods.