South America
A Methodology for Creating Question Answering Corpora Using Inverse Data Annotation
Deriu, Jan, Mlynchyk, Katsiaryna, Schläpfer, Philippe, Rodrigo, Alvaro, von Grünigen, Dirk, Kaiser, Nicolas, Stockinger, Kurt, Agirre, Eneko, Cieliebak, Mark
In this paper, we introduce a novel methodology to efficiently construct a corpus for question answering over structured data. For this, we introduce an intermediate representation that is based on the logical query plan in a database called Operation Trees (OT). This representation allows us to invert the annotation process without losing flexibility in the types of queries that we generate. Furthermore, it allows for fine-grained alignment of query tokens to OT operations. In our method, we randomly generate OTs from a context-free grammar. Afterwards, annotators have to write the appropriate natural language question that is represented by the OT. Finally, the annotators assign the tokens to the OT operations. We apply the method to create a new corpus OTTA (Operation Trees and Token Assignment), a large semantic parsing corpus for evaluating natural language interfaces to databases. We compare OTTA to Spider and LC-QuaD 2.0 and show that our methodology more than triples the annotation speed while maintaining the complexity of the queries. Finally, we train a state-of-the-art semantic parsing model on our data and show that our corpus is a challenging dataset and that the token alignment can be leveraged to increase the performance significantly.
Israeli Innovators Harness Artificial Intelligence Technologies To Curb The Global COVID-19 Pandemic
As the number of people who've tested positive for coronavirus is mounting and could reach 2 million in the coming days, Israeli innovators are harnessing artificial intelligence technologies to curb the global pandemic, perhaps the most challenging public health crisis in modern history. What we know already is that scientists and researchers are working diligently to find treatments and to develop a vaccine for coronavirus. Meanwhile, artificial intelligence technologies are emerging as key solutions to combatting coronavirus, and Israel is well positioned in this field. Israel is well known for its strength in deep-tech, and is also home to a vibrant AI ecosystem that has been growing rapidly over the past few years. Israel's unique tech ecosystem includes companies and startups that utilize AI technologies in healthcare, cybersecurity, autonomous driving, and many other fields.
Google apologizes after its Vision AI produced racist results
A Google service that automatically labels images produced starkly different results depending on skin tone on a given image. The company fixed the issue, but the problem is likely much broader. In the fight against the novel coronavirus, many countries ordered that citizens have their temperature checked at train stations or airports. The device needed in such situations, a hand-held thermometer, has risen from a specialist item to a common sight. A branch of Artificial Intelligence known as "computer vision" focuses on automated image labeling.
Flattening the curves: on-off lock-down strategies for COVID-19 with an application to Brazi
Tarrataca, L., Dias, C. M., Haddad, D. B., Arruda, E. F.
The current COVID-19 pandemic is affecting different countries in different ways. The assortment of reporting techniques alongside other issues, such as underreporting and budgetary constraints, makes predicting the spread and lethality of the virus a challenging task. This work attempts to gain a better understanding of how COVID-19 will affect one of the least studied countries, namely Brazil. Currently, several Brazilian states are in a state of lock-down. However, there is political pressure for this type of measures to be lifted. This work considers the impact that such a termination would have on how the virus evolves locally. This was done by extending the SEIR model with an on / off strategy. Given the simplicity of SEIR we also attempted to gain more insight by developing a neural regressor. We chose to employ features that current clinical studies have pinpointed has having a connection to the lethality of COVID-19. We discuss how this data can be processed in order to obtain a robust assessment.
Diverse Instances-Weighting Ensemble based on Region Drift Disagreement for Concept Drift Adaptation
Liu, Anjin, Lu, Jie, Zhang, Guangquan
Concept drift refers to changes in the distribution of underlying data and is an inherent property of evolving data streams. Ensemble learning, with dynamic classifiers, has proved to be an efficient method of handling concept drift. However, the best way to create and maintain ensemble diversity with evolving streams is still a challenging problem. In contrast to estimating diversity via inputs, outputs, or classifier parameters, we propose a diversity measurement based on whether the ensemble members agree on the probability of a regional distribution change. In our method, estimations over regional distribution changes are used as instance weights. Constructing different region sets through different schemes will lead to different drift estimation results, thereby creating diversity. The classifiers that disagree the most are selected to maximize diversity. Accordingly, an instance-based ensemble learning algorithm, called the diverse instance weighting ensemble (DiwE), is developed to address concept drift for data stream classification problems. Evaluations of various synthetic and real-world data stream benchmarks show the effectiveness and advantages of the proposed algorithm.
Anomaly Detection in Trajectory Data with Normalizing Flows
Dias, Madson L. D., Mattos, César Lincoln C., da Silva, Ticiana L. C., de Macedo, José Antônio F., Silva, Wellington C. P.
The task of detecting anomalous data patterns is as important in practical applications as challenging. In the context of spatial data, recognition of unexpected trajectories brings additional difficulties, such as high dimensionality and varying pattern lengths. We aim to tackle such a problem from a probability density estimation point of view, since it provides an unsupervised procedure to identify out of distribution samples. More specifically, we pursue an approach based on normalizing flows, a recent framework that enables complex density estimation from data with neural networks. Our proposal computes exact model likelihood values, an important feature of normalizing flows, for each segment of the trajectory. Then, we aggregate the segments' likelihoods into a single coherent trajectory anomaly score. Such a strategy enables handling possibly large sequences with different lengths. We evaluate our methodology, named aggregated anomaly detection with normalizing flows (GRADINGS), using real world trajectory data and compare it with more traditional anomaly detection techniques. The promising results obtained in the performed computational experiments indicate the feasibility of the GRADINGS, specially the variant that considers autoregressive normalizing flows.
Quantifying Notes Revisited
To a multi-agent logic of knowledge or belief we can add public announcements to model publicly observed information change, or action models to model information change that is differently observed by different agents, but also modalities representing quantification over such information change, such as quantifiers over announcements or quantifiers over actions models. Such additions may result in more complex or undecidable logics, and create a very open landscape of relative expressivity. The survey [88] of such logics focused on open problems. Some such open problems have since then been resolved, and yet others have come to the fore. In this updated survey we review what is known about such logics with quantification over information change, including digressions into what are known as relation changing modal(but often not epistemic) logics. Again we focus on open problems.
AI (Artificial Intelligence) Projects: Where To Start?
Artificial Intelligence (AI) is clearly a must-have when it comes to being competitive in today's markets. But implementing this technology has been challenging, even for some of the world's top companies. There are issues with data, finding the right talent and creating models that generate sufficient ROI. As a result, many AI projects fail. According to IDC, only abut 35% of organizations succeed in getting models into production successfully.
Attribute-based Regularization of VAE Latent Spaces
Selective manipulation of data attributes using deep generative models is an active area of research. In this paper, we present a novel method to structure the latent space of a Variational Auto-Encoder (VAE) to encode different continuous-valued attributes explicitly. This is accomplished by using an attribute regularization loss which enforces a monotonic relationship between the attribute values and the latent code of the dimension along which the attribute is to be encoded. Consequently, post-training, the model can be used to manipulate the attribute by simply changing the latent code of the corresponding regularized dimension. The results obtained from several quantitative and qualitative experiments show that the proposed method leads to disentangled and interpretable latent spaces that can be used to effectively manipulate a wide range of data attributes spanning image and symbolic music domains.
Training Data Set Assessment for Decision-Making in a Multiagent Landmine Detection Platform
Florez-Lozano, Johana, Caraffini, Fabio, Parra, Carlos, Gongora, Mario
Real-world problems such as landmine detection require multiple sources of information to reduce the uncertainty of decision-making. A novel approach to solve these problems includes distributed systems, as presented in this work based on hardware and software multi-agent systems. To achieve a high rate of landmine detection, we evaluate the performance of a trained system over the distribution of samples between training and validation sets. Additionally, a general explanation of the data set is provided, presenting the samples gathered by a cooperative multi-agent system developed for detecting improvised explosive devices. The results show that input samples affect the performance of the output decisions, and a decision-making system can be less sensitive to sensor noise with intelligent systems obtained from a diverse and suitably organised training set.