Goto

Collaborating Authors

 South America


On the optimality of kernels for high-dimensional clustering

arXiv.org Machine Learning

This paper studies the optimality of kernel methods in high-dimensional data clustering. Recent works have studied the large sample performance of kernel clustering in the high-dimensional regime, where Euclidean distance becomes less informative. However, it is unknown whether popular methods, such as kernel k-means, are optimal in this regime. We consider the problem of high-dimensional Gaussian clustering and show that, with the exponential kernel function, the sufficient conditions for partial recovery of clusters using the NP-hard kernel k-means objective matches the known information-theoretic limit up to a factor of $\sqrt{2}$ for large $k$. It also exactly matches the known upper bounds for the non-kernel setting. We also show that a semi-definite relaxation of the kernel k-means procedure matches up to constant factors, the spectral threshold, below which no polynomial-time algorithm is known to succeed. This is the first work that provides such optimality guarantees for the kernel k-means as well as its convex relaxation. Our proofs demonstrate the utility of the less known polynomial concentration results for random variables with exponentially decaying tails in a higher-order analysis of kernel methods.


Conformance Checking Approximation using Subset Selection and Edit Distance

arXiv.org Artificial Intelligence

Conformance checking techniques let us find out to what degree a process model and real execution data correspond to each other. In recent years, alignments have proven extremely useful in calculating conformance statistics. Most techniques to compute alignments provide an exact solution. However, in many applications, it is enough to have an approximation of the conformance value. Specifically, for large event data, the computing time for alignments is considerably long using current techniques which makes them inapplicable in reality. Also, it is no longer feasible to use standard hardware for complex processes. Hence, we need techniques that enable us to obtain fast, and at the same time, accurate approximation of the conformance values. This paper proposes new approximation techniques to compute approximated conformance checking values close to exact solution values in a faster time. Those methods also provide upper and lower bounds for the approximated alignment value. Our experiments on real event data show that it is possible to improve the performance of conformance checking by using the proposed methods compared to using the state-of-the-art alignment approximation technique. Results show that in most of the cases, we provide tight bounds, accurate approximated alignment values, and similar deviation statistics.


Knowledge Infused Learning (K-IL): Towards Deep Incorporation of Knowledge in Deep Learning

arXiv.org Artificial Intelligence

Learning the underlying patterns in the data goes beyond instance-based generalization to some external knowledge represented in structured graphs or networks. Deep Learning (DL) has shown significant advances in probabilistically learning latent patterns in the data using a multi-layered network of computational nodes (i.e. neurons/hidden units). However, with the tremendous amount of training data, uncertainty in generalization on domain-specific tasks, and delta improvement with an increase in complexity of models seem to raise a concern on the features learned by the model. As incorporation of domain specific knowledge will aid in supervising the learning of features for the model, infusion of knowledge from knowledge graphs within hidden layers will further enhance the learning process. Although much work remains, we believe that KGs will play an increasing role in developing hybrid neuro-symbolic intelligent systems (that is bottom up deep learning with top down symbolic computing) as well as in building explainable AI systems for which KGs will provide a scaffolding for punctuating neural computing. In this position paper, we describe our motivation for such hybrid approach and a framework that combines knowledge graph and neural networks.


My Journey South: Tracing developments on Artificial Intelligence (AI) in Latin America and the Caribbean – SRC

#artificialintelligence

While some still consider AI to be beyond the grasp of developing countries, our South American neighbours have been shattering that stereotype. AI is being deployed in a number of their endeavours: to speed up artefact findings in Peru; to increase crop yields in Colombian rice fields through AI-powered platforms; to boost security and enhance customer service in Brazil's banking sector; to create vegan alternatives with the same taste and texture as animal-based foods in Chile's food industry; to predict school dropouts and teenage pregnancy in Argentina; and to forecast crimes in Uruguay. Some of the push in AI adoption in these countries has come from academics and researchers, like the ones at the University of Sao Paulo who are developing AI to determine the susceptibility of patients to disease outbreaks; or Peru's National Engineering University where robots are being used for mine exploration to detect gases; or Argentina's National Scientific and Technical Research Council where AI software is predicting early onset pluripotent stem cell differentiation. These and other truths were revealed to me at a Latin America and Caribbean (LAC) Workshop on AI organized by Facebook and the Inter-American Development Bank in Montevideo, Uruguay, in November this year. I was the lone Caribbean participant in attendance, presenting my paper entitled: AI & The Caribbean: A Discussion on Potential Applications & Ethical Considerations, on behalf of the Shridath Ramphal Centre (UWI, Cave Hill).


A Surrogate Video-Based Safety Methodology for Diagnosis and Evaluation of Low-Cost Pedestrian-Safety Countermeasures: The Case of Cochabamba, Bolivia

#artificialintelligence

Due to a lack of reliable data collection systems, traffic fatalities and injuries are often under-reported in developing countries. Recent developments in surrogate road safety methods and video analytics tools offer an alternative approach that can be both lower cost and more time efficient when crash data is incomplete or missing. However, very few studies investigating pedestrian road safety in developing countries using these approaches exist. This research uses an automated video analytics tool to develop and analyze surrogate traffic safety measures and to evaluate the effectiveness of temporary low-cost countermeasures at selected pedestrian crossings at risky intersections in the city of Cochabamba, Bolivia. Specialized computer vision software is used to process hundreds of hours of video data and generate data on road users' speed and trajectories.


futureofwork _2019-11-26_19-00-43.xlsx

#artificialintelligence

The graph represents a network of 3,989 Twitter users whose tweets in the requested range contained "futureofwork ", or who were replied to or mentioned in those tweets. The network was obtained from the NodeXL Graph Server on Wednesday, 27 November 2019 at 03:02 UTC. The requested start date was Monday, 25 November 2019 at 01:01 UTC and the maximum number of days (going backward) was 14. The maximum number of tweets collected was 5,000. The tweets in the network were tweeted over the 3-day, 1-hour, 59-minute period from Thursday, 21 November 2019 at 23:00 UTC to Monday, 25 November 2019 at 01:00 UTC.


AI has a bias problem. Barring African experts from a conference in Canada won't help

#artificialintelligence

London (CNN Business)Some of the leading artificial intelligence experts from Africa and South America have been denied visas to attend a major industry conference in Canada, dealing a setback to efforts to prevent bias from taking root in the new technology. Conference organizers say Canadian immigration authorities have denied visas to two dozen academics from countries such as Nigeria and Brazil, preventing them from attending the event next month in Vancouver. Katherine Heller, a professor who serves as co-chair of diversity and inclusion at the Neural Information Processing Systems conference, said organizers "are trying extremely hard" to have the visa denials overturned. "It is very significant for the field of AI that all voices be heard," she said. The problem of algorithmic bias in data science has become more pronounced, and there's mounting evidence that AI-powered algorithms display bias against women and some racial groups.


AI magic bean could save farmers millions

#artificialintelligence

Farmers across the world could jack up giant profits using an Artificial Intelligence soil monitoring system developed at Brunel University London. By collecting data about soil and growing conditions, the'magic bean' helps farmers boost crops, cut waste and save time, money and water. It comes after France this year saw record temperatures of 49.5 ºC, the US had its wettest spring since 1995 and severe frost threatened Brazil's coffee harvest. The Brunel algorithms could help producers work around freak weather triggered by climate change and unplanned supply problems after Brexit. "We have a way of using data to make crops grow better, worldwide," said electronic engineer Dr Tatiana Kalganova.


Financial Time Series Forecasting with Deep Learning : A Systematic Literature Review: 2005-2019

arXiv.org Machine Learning

Financial time series forecasting is, without a doubt, the top choice of computational intelligence for finance researchers from both academia and financial industry due to its broad implementation areas and substantial impact. Machine Learning (ML) researchers came up with various models and a vast number of studies have been published accordingly. As such, a significant amount of surveys exist covering ML for financial time series forecasting studies. Lately, Deep Learning (DL) models started appearing within the field, with results that significantly outperform traditional ML counterparts. Even though there is a growing interest in developing models for financial time series forecasting research, there is a lack of review papers that were solely focused on DL for finance. Hence, our motivation in this paper is to provide a comprehensive literature review on DL studies for financial time series forecasting implementations. We not only categorized the studies according to their intended forecasting implementation areas, such as index, forex, commodity forecasting, but also grouped them based on their DL model choices, such as Convolutional Neural Networks (CNNs), Deep Belief Networks (DBNs), Long-Short Term Memory (LSTM). We also tried to envision the future for the field by highlighting the possible setbacks and opportunities, so the interested researchers can benefit.


Applications of the Deep Galerkin Method to Solving Partial Integro-Differential and Hamilton-Jacobi-Bellman Equations

arXiv.org Machine Learning

We extend the Deep Galerkin Method (DGM) introduced in Sirignano and Spiliopoulos (2018) to solve a number of partial differential equations (PDEs) that arise in the context of optimal stochastic control and mean field games. First, we consider PDEs where the function is constrained to be positive and integrate to unity, as is the case with Fokker-Planck equations. Our approach involves reparameterizing the solution as the exponential of a neural network appropriately normalized to ensure both requirements are satisfied. This then gives rise to a partial integro-differential equation (PIDE) where the integral appearing in the equation is handled using importance sampling. Secondly, we tackle a number of Hamilton-Jacobi-Bellman (HJB) equations that appear in stochastic optimal control problems. The key contribution is that these equations are approached in their unsimplified primal form which includes an optimization problem as part of the equation. We extend the DGM algorithm to solve for the value function and the optimal control simultaneously by characterizing both as deep neural networks. Training the networks is performed by taking alternating stochastic gradient descent steps for the two functions, a technique similar in spirit to policy improvement algorithms.