Oceania
A Benchmarking Study of Embedding-based Entity Alignment for Knowledge Graphs
Sun, Zequn, Zhang, Qingheng, Hu, Wei, Wang, Chengming, Chen, Muhao, Akrami, Farahnaz, Li, Chengkai
Entity alignment seeks to find entities in different knowledge graphs (KGs) that refer to the same real-world object. Recent advancement in KG embedding impels the advent of embedding-based entity alignment, which encodes entities in a continuous embedding space and measures entity similarities based on the learned embeddings. In this paper, we conduct a comprehensive experimental study of this emerging field. This study surveys 23 recent embedding-based entity alignment approaches and categorizes them based on their techniques and characteristics. We further observe that current approaches use different datasets in evaluation, and the degree distributions of entities in these datasets are inconsistent with real KGs. Hence, we propose a new KG sampling algorithm, with which we generate a set of dedicated benchmark datasets with various heterogeneity and distributions for a realistic evaluation. This study also produces an open-source library, which includes 12 representative embedding-based entity alignment approaches. We extensively evaluate these approaches on the generated datasets, to understand their strengths and limitations. Additionally, for several directions that have not been explored in current approaches, we perform exploratory experiments and report our preliminary findings for future studies. The benchmark datasets, open-source library and experimental results are all accessible online and will be duly maintained.
Fair Allocation with Diminishing Differences
Segal-Halevi, Erel | Hassidim, Avinatan (Bar-Ilan University) | Aziz, Haris (UNSW Sydney and Data61 CSIRO)
Ranking alternatives is a natural way for humans to explain their preferences. It is used in many settings, such as school choice, course allocations and residency matches. Without having any information on the underlying cardinal utilities, arguing about the fairness of allocations requires extending the ordinal item ranking to ordinal bundle ranking. The most commonly used such extension is stochastic dominance (SD), where a bundle X is preferred over a bundle Y if its score is better according to all additive score functions. SD is a very conservative extension, by which few allocations are necessarily fair while many allocations are possibly fair. We propose to make a natural assumption on the underlying cardinal utilities of the players, namely that the difference between two items at the top is larger than the difference between two items at the bottom. This assumption implies a preference extension which we call diminishing differences (DD), where X is preferred over Y if its score is better according to all additive score functions satisfying the DD assumption. We give a full characterization of allocations that are necessarily-proportional or possibly-proportional according to this assumption. Based on this characterization, we present a polynomial-time algorithm for finding a necessarily-DD-proportional allocation whenever it exists. Using simulations, we compare the various fairness criteria in terms of their probability of existence, and their probability of being fair by the underlying cardinal valuations. We find that necessary-DD-proportionality fares well in both measures. We also consider envy-freeness and Pareto optimality under diminishing-differences, as well as chore allocation under the analogous condition --- increasing-differences.
Multivariate Functional Regression via Nested Reduced-Rank Regularization
Liu, Xiaokang, Ma, Shujie, Chen, Kun
We propose a nested reduced-rank regression (NRRR) approach in fitting regression model with multivariate functional responses and predictors, to achieve tailored dimension reduction and facilitate interpretation/visualization of the resulting functional model. Our approach is based on a two-level low-rank structure imposed on the functional regression surfaces. A global low-rank structure identifies a small set of latent principal functional responses and predictors that drives the underlying regression association. A local low-rank structure then controls the complexity and smoothness of the association between the principal functional responses and predictors. Through a basis expansion approach, the functional problem boils down to an interesting integrated matrix approximation task, where the blocks or submatrices of an integrated low-rank matrix share some common row space and/or column space. An iterative algorithm with convergence guarantee is developed. We establish the consistency of NRRR and also show through non-asymptotic analysis that it can achieve at least a comparable error rate to that of the reduced-rank regression. Simulation studies demonstrate the effectiveness of NRRR. We apply NRRR in an electricity demand problem, to relate the trajectories of the daily electricity consumption with those of the daily temperatures.
Multi-Objective Variational Autoencoder: an Application for Smart Infrastructure Maintenance
Anaissi, Ali, Zandavi, Seid Miad
Multi-way data analysis has become an essential tool for capturing underlying structures in higher-order data sets where standard two-way analysis techniques often fail to discover the hidden correlations between variables in multi-way data. We propose a multi-objective variational autoencoder (MVA) method for smart infrastructure damage detection and diagnosis in multi-way sensing data based on the reconstruction probability of autoencoder deep neural network (ADNN). Our method fuses data from multiple sensors in one ADNN at which informative features are being extracted and utilized for damage identification. It generates probabilistic anomaly scores to detect damage, asses its severity and further localize it via a new localization layer introduced in the ADNN. We evaluated our method on multi-way datasets in the area of structural health monitoring for damage diagnosis purposes. The data was collected from our deployed data acquisition system on a cable-stayed bridge in Western Sydney and from a laboratory based building structure obtained from Los Alamos National Laboratory (LANL). Experimental results show that the proposed method can accurately detect structural damage. It was also able to estimate the different levels of damage severity, and capture damage locations in an unsupervised aspect. Compared to the state-of-the-art approaches, our proposed method shows better performance in terms of damage detection and localization.
KALE: When Energy-Based Learning Meets Adversarial Training
Arbel, Michael, Zhou, Liang, Gretton, Arthur
Legendre duality provides a variational lower-bound for the Kullback-Leibler divergence (KL) which can be estimated using samples, without explicit knowledge of the density ratio. We use this estimator, the \textit{KL Approximate Lower-bound Estimate} (KALE), in a contrastive setting for learning energy-based models, and show that it provides a maximum likelihood estimate (MLE). We then extend this procedure to adversarial training, where the discriminator represents the energy and the generator is the base measure of the energy-based model. Unlike in standard generative adversarial networks (GANs), the learned model makes use of both generator and discriminator to generate samples. This is achieved using Hamiltonian Monte Carlo in the latent space of the generator, using information from the discriminator, to find regions in that space that produce better quality samples. We also show that, unlike the KL, KALE enjoys smoothness properties that make it suitable for adversarial training, and provide convergence rates for KALE when the negative log density ratio belongs to the variational family. Finally, we demonstrate the effectiveness of this approach on simple datasets.
Channel Attention with Embedding Gaussian Process: A Probabilistic Methodology
Xie, Jiyang, Chang, Dongliang, Ma, Zhanyu, Zhang, Guoqiang, Guo, Jun
Channel attention mechanisms, as the key components of some modern convolutional neural networks (CNNs) architectures, have been commonly used in many visual tasks for effective performance improvement. It is able to reinforce the informative channels and to suppress useless channels of feature maps obtained by CNNs. Recently, different attention modules have been proposed, which are implemented in various ways. However, they are mainly based on convolution and pooling operations, which are lack of intuitive and reasonable insights about the principles that they are based on. Moreover, the ways that they improve the performance of the CNNs is not clear either. In this paper, we propose a Gaussian process embedded channel attention (GPCA) module and interpret the channel attention intuitively and reasonably in a probabilistic way. The GPCA module is able to model the correlations from channels which are assumed as beta distributed variables with Gaussian process prior. As the beta distribution is intractably integrated into the end-to-end training of the CNNs, we utilize an appropriate approximation of the beta distribution to make the distribution assumption implemented easily. In this case, the proposed GPCA module can be integrated into the end-to-end training of the CNNs. Experimental results demonstrate that the proposed GPCA module can improve the accuracies of image classification on four widely used datasets.
2019 AI Index Report: R&D in AI Continues to Increase - EnterpriseTalk
The US is a leader in investing capital into private AI with nearly US$12 billion. China, which came second with US$6.8 billion investment, also files more AI patents than any other country across the globe and three times more than Japan. The majority of AI patents filed between 2014-2018 were filed in the U.S. and Canada, and 94% of patents are filed in wealthy nations. Mergers and acquisitions worth $37 billion were spurred thanks to AI. At the same time, IPOs worth $34 billion were also associated with AI. Investment in AI startups recorded a rapid increase in the last ten years from a total of $1.3 billion raised in 2010 to over $40.4 billion.
Enlisting analytics and AI to contain the next pandemic
Much has been written about the coronavirus since it was first identified in China in January and much more will undoubtedly be written before the subsequently alarming spread abates and medical science comes up with an effective cure. And while news of the steady increase in reported numbers of people infected by and dying from COVID-19, as it is now known, has been dire, the good news is that we are getting much better at predicting and tracking the spread of infectious diseases. Three out of four infectious diseases originate in other species but their rapid spread in humans is facilitated by our ever-increasing mobility. International travel is now such that a disease that might once have stayed relatively contained can now spread across the world in mere weeks. We saw this with the Severe Acute Respiratory Syndrome (SARS) virus in 2003 and we see it again today.
Enlisting analytics and AI to contain the next pandemic
Much has been written about the coronavirus since it was first identified in China in January and much more will undoubtedly be written before the subsequently alarming spread abates and medical science comes up with an effective cure. And while news of the steady increase in reported numbers of people infected by and dying from COVID-19, as it is now known, has been dire, the good news is that we are getting much better at predicting and tracking the spread of infectious diseases. Three out of four infectious diseases originate in other species but their rapid spread in humans is facilitated by our ever-increasing mobility. International travel is now such that a disease that might once have stayed relatively contained can now spread across the world in mere weeks. We saw this with the Severe Acute Respiratory Syndrome (SARS) virus in 2003 and we see it again today.
MD talks Artificial Intelligence and insurance
"We're a world leading artificial intelligence (AI) platform that monitors real time and real-world events and provides all-round risk detection and solutions," Rod Moynihan told Insurance Business. Moynihan is the managing director of Dataminr in Australia and New Zealand, a global real-time information discovery company that is pioneering what it sees as ground-breaking technology for detecting, classifying, and determining the significance of public information in real time. Moynihan recently gave an interview to Insurance Business to explain how the platform works and how it can aid insurers. Using public information and data available from across the world, Dataminr's AI platform finds, dissects and quantifies a large amount of data to make sense of potentially large impact events that can affect customers. "It sorts through millions upon millions of publicly available data," explained Moynihan.