Goto

Collaborating Authors

 Europe


An experimental study of graph-based semi-supervised classification with additional node information

arXiv.org Machine Learning

The volume of data generated by internet and social networks is increasing every day, and there is a clear need for efficient ways of extracting useful information from them. As those data can take different forms, it is important to use all the available data representations for prediction. In this paper, we focus our attention on supervised classification using both regular plain, tabular, data and structural information coming from a network structure. 14 techniques are investigated and compared in this study and can be divided in three classes: the first one uses only the plain data to build a classification model, the second uses only the graph structure and the last uses both information sources. The relative performances in these three cases are investigated. Furthermore, the effect of using a graph embedding and well-known indicators in spatial statistics is also studied. Possible applications are automatic classification of web pages or other linked documents, of people in a social network or of proteins in a biological complex system, to name a few. Based on our comparison, we draw some general conclusions and advices to tackle this particular classification task: some datasets can be better explained by their graph structure (graph-driven), or by their feature set (features-driven). The most efficient methods are discussed in both cases.


Non-Stationary Spectral Kernels

arXiv.org Machine Learning

We propose non-stationary spectral kernels for Gaussian process regression. We propose to model the spectral density of a non-stationary kernel function as a mixture of input-dependent Gaussian process frequency density surfaces. We solve the generalised Fourier transform with such a model, and present a family of non-stationary and non-monotonic kernels that can learn input-dependent and potentially long-range, non-monotonic covariances between inputs. We derive efficient inference using model whitening and marginalized posterior, and show with case studies that these kernels are necessary when modelling even rather simple time series, image or geospatial data with non-stationary characteristics.


Power Systems Data Fusion based on Belief Propagation

arXiv.org Machine Learning

The increasing complexity of the power grid, due to higher penetration of distributed resources and the growing availability of interconnected, distributed metering devices re- quires novel tools for providing a unified and consistent view of the system. A computational framework for power systems data fusion, based on probabilistic graphical models, capable of combining heterogeneous data sources with classical state estimation nodes and other customised computational nodes, is proposed. The framework allows flexible extension of the notion of grid state beyond the view of flows and injection in bus-branch models, and an efficient, naturally distributed inference algorithm can be derived. An application of the data fusion model to the quantification of distributed solar energy is proposed through numerical examples based on semi-synthetic simulations of the standard IEEE 14-bus test case.


Anti-spoofing Methods for Automatic SpeakerVerification System

arXiv.org Machine Learning

Growing interest in automatic speaker verification (ASV)systems has lead to significant quality improvement of spoofing attackson them. Many research works confirm that despite the low equal er-ror rate (EER) ASV systems are still vulnerable to spoofing attacks. Inthis work we overview different acoustic feature spaces and classifiersto determine reliable and robust countermeasures against spoofing at-tacks. We compared several spoofing detection systems, presented so far,on the development and evaluation datasets of the Automatic SpeakerVerification Spoofing and Countermeasures (ASVspoof) Challenge 2015.Experimental results presented in this paper demonstrate that the useof magnitude and phase information combination provides a substantialinput into the efficiency of the spoofing detection systems. Also wavelet-based features show impressive results in terms of equal error rate. Inour overview we compare spoofing performance for systems based on dif-ferent classifiers. Comparison results demonstrate that the linear SVMclassifier outperforms the conventional GMM approach. However, manyresearchers inspired by the great success of deep neural networks (DNN)approaches in the automatic speech recognition, applied DNN in thespoofing detection task and obtained quite low EER for known and un-known type of spoofing attacks.


Boundary Crossing Probabilities for General Exponential Families

arXiv.org Machine Learning

We consider parametric exponential families of dimension $K$ on the real line. We study a variant of \textit{boundary crossing probabilities} coming from the multi-armed bandit literature, in the case when the real-valued distributions form an exponential family of dimension $K$. Formally, our result is a concentration inequality that bounds the probability that $\mathcal{B}^\psi(\hat \theta_n,\theta^\star)\geq f(t/n)/n$, where $\theta^\star$ is the parameter of an unknown target distribution, $\hat \theta_n$ is the empirical parameter estimate built from $n$ observations, $\psi$ is the log-partition function of the exponential family and $\mathcal{B}^\psi$ is the corresponding Bregman divergence. From the perspective of stochastic multi-armed bandits, we pay special attention to the case when the boundary function $f$ is logarithmic, as it is enables to analyze the regret of the state-of-the-art \KLUCB\ and \KLUCBp\ strategies, whose analysis was left open in such generality. Indeed, previous results only hold for the case when $K=1$, while we provide results for arbitrary finite dimension $K$, thus considerably extending the existing results. Perhaps surprisingly, we highlight that the proof techniques to achieve these strong results already existed three decades ago in the work of T.L. Lai, and were apparently forgotten in the bandit community. We provide a modern rewriting of these beautiful techniques that we believe are useful beyond the application to stochastic multi-armed bandits.


Audio-replay attack detection countermeasures

arXiv.org Machine Learning

This paper presents the Speech Technology Center (STC) replay attack detection systems proposed for Automatic Speaker Verification Spoofing and Countermeasures Challenge 2017. In this study we focused on comparison of different spoofing detection approaches. These were GMM based methods, high level features extraction with simple classifier and deep learning frameworks. Experiments performed on the development and evaluation parts of the challenge dataset demonstrated stable efficiency of deep learning approaches in case of changing acoustic conditions. At the same time SVM classifier with high level features provided a substantial input in the efficiency of the resulting STC systems according to the fusion systems results.


Artificial Intelligence Is The Buzzword Of 2017, Says Creathor Venture

#artificialintelligence

How would you describe Creathor Venture in a few words? Creathor is a pan-European venture fund based in Germany, Switzerland, and Sweden. We have been successfully investing in early-stage technology-oriented companies and entrepreneurs for more than 30 years. We invest up to EUR 10 million per startup over several rounds. We currently manage funds of more than EUR 220 million.


Half of World's Languages Could Be Extinct by 2100

U.S. News

But modern tools are helping to revive Ireland's national language. An Irish proverb advises that it is often wise for one to hold his tongue. An tรฉ is ciรบine is รฉ is buaine, or "he who is silent is the stronger." But that ancestral wisdom isn't the best policy when the very language it comes from is threatened. The Irish language, Gaelic, is one of more than 40 percent of the world's 6,000 spoken languages that are endangered, according to UNESCO.


Azure Craft A hack for kids and parents to learn AI Programming

#artificialintelligence

AzureCraft was conceived in 2016 by myself, Richard Conway, the co-founder of the UK Azure Users Group, Andy Cross and Allan Mitchell, one of the founders of SQLBITS, the largest SQL conference in Europe. Our idea stemmed from the fact that our children were either continually playing Minecraft or other games such as Roblox which looked a lot like Minecraft. After speaking to a lot of parents at schools I realised that there was little room in parents minds to see Minecraft as a creative canvas. Most saw it as a waste of time and others were irate at the parade of foul-mouthed teens that their kids watched on YouTube that spoke about Minecraft. Most parents were unaware that their children were using Minecraft and scratch at schools to automate the creation the of their worlds.


A Knowledge Graph-based Semantic Database for Biomedical Sciences

@machinelearnbot

In this article, we talk with one of our users: Antonio Messina from the High Performance Computing and Networking Institute of the Italian National Research Council (ICAR-CNR). Antonio (@xMAnton on Twitter) is a Computer Science Engineer who works as an Applied Scientist at the largest public research institution in Italy. His area of expertise includes (No)SQL databases and advanced Unix systems administration, and he likes to get his hands dirty coding mainly in Java and Node.js. He is enthusiastic about technologies such as graph databases and Docker and is constantly looking for innovation in IT. Recently, Antonio successfully submitted a paper that describes a practical use case for GRAKN.AI.