Goto

Collaborating Authors

 South America


Double machine learning for sample selection models

arXiv.org Machine Learning

This paper considers treatment evaluation when outcomes are only observed for a subpopulation due to sample selection or outcome attrition/non-response. For identification, we combine a selection-on-observables assumption for treatment assignment with either selection-on-observables or instrumental variable assumptions concerning the outcome attrition/sample selection process. To control in a data-driven way for potentially high dimensional pre-treatment covariates that motivate the selectionon-observables assumptions, we adapt the double machine learning framework to sample selection problems. That is, we make use of (a) Neyman-orthogonal and doubly robust score functions, which imply the robustness of treatment effect estimation to moderate regularization biases in the machine learningbased estimation of the outcome, treatment, or sample selection models and (b) sample splitting (or cross-fitting) to prevent overfitting bias. We demonstrate that the proposed estimators are asymptotically normal and root-n consistent under specific regularity conditions concerning the machine learners and investigate their finite sample properties in a simulation study. The estimator is available in the causalweight package for the statistical software R. Keywords: sample selection, double machine learning, doubly robust estimation, efficient score.


Archaeology: Ancient Amazons laid out their villages like a clock face to represent the cosmos

Daily Mail - Science & tech

Ancient Amazonians laid out their settlements in circles 700 years ago -- with radiating mounds and roads as may have represented the cosmos -- a study found. Experts led from Exeter used lidar-based sensing equipment mounted on helicopters to see below the canopy of the overlying rainforest in south Acre State, Brazil. The 35 mounded villages were constructed to the distinctive and repeated pattern by the ancient Acreans between around 1300–1700 AD. Deforestation and archaeological digs in Acre State has previously revealed the presence of large earthworks and circular mound villages. However, the full extent of the constructions, their layouts and their organisation across the region had been obscured by the dense forest until now.


Forecasting the Olympic medal distribution during a pandemic: a socio-economic machine learning model

arXiv.org Machine Learning

Forecasting the number of Olympic medals for each nation is highly relevant for different stakeholders: Ex ante, sports betting companies can determine the odds while sponsors and media companies can allocate their resources to promising teams. Ex post, sports politicians and managers can benchmark the performance of their teams and evaluate the drivers of success. To significantly increase the Olympic medal forecasting accuracy, we apply machine learning, more specifically a two-staged Random Forest, thus outperforming more traditional na\"ive forecast for three previous Olympics held between 2008 and 2016 for the first time. Regarding the Tokyo 2020 Games in 2021, our model suggests that the United States will lead the Olympic medal table, winning 120 medals, followed by China (87) and Great Britain (74). Intriguingly, we predict that the current COVID-19 pandemic will not significantly alter the medal count as all countries suffer from the pandemic to some extent (data inherent) and limited historical data points on comparable diseases (model inherent).


Cross-Modal Collaborative Representation Learning and a Large-Scale RGBT Benchmark for Crowd Counting

arXiv.org Artificial Intelligence

Crowd counting is a fundamental yet challenging problem, which desires rich information to generate pixel-wise crowd density maps. However, most previous methods only utilized the limited information of RGB images and may fail to discover the potential pedestrians in unconstrained environments. In this work, we find that incorporating optical and thermal information can greatly help to recognize pedestrians. To promote future researches in this field, we introduce a large-scale RGBT Crowd Counting (RGBT-CC) benchmark, which contains 2,030 pairs of RGB-thermal images with 138,389 annotated people. Furthermore, to facilitate the multimodal crowd counting, we propose a cross-modal collaborative representation learning framework, which consists of multiple modality-specific branches, a modality-shared branch, and an Information Aggregation-Distribution Module (IADM) to fully capture the complementary information of different modalities. Specifically, our IADM incorporates two collaborative information transfer components to dynamically enhance the modality-shared and modality-specific representations with a dual information propagation mechanism. Extensive experiments conducted on the RGBT-CC benchmark demonstrate the effectiveness of our framework for RGBT crowd counting. Moreover, the proposed approach is universal for multimodal crowd counting and is also capable to achieve superior performance on the ShanghaiTechRGBD dataset.


River: machine learning for streaming data in Python

arXiv.org Artificial Intelligence

River is a machine learning library for dynamic data streams and continual learning. It provides multiple state-of-the-art learning methods, data generators/transformers, performance metrics and evaluators for different stream learning problems. It is the result from the merger of the two most popular packages for stream learning in Python: Creme and scikit-multiflow. River introduces a revamped architecture based on the lessons learnt from the seminal packages. River's ambition is to be the go-to library for doing machine learning on streaming data. Additionally, this open source package brings under the same umbrella a large community of practitioners and researchers. The source code is available at https://github.com/online-ml/river.


RPA - 10 Powerful Examples in Enterprise - Algorithm-X Lab

#artificialintelligence

More and more enterprises are turning to a promising technology called RPA (robotic process automation) to become more productive and efficient. Successful implementation also helps to cut costs and reduce error rates. RPA can automate mundane and predictable tasks and processes leaving employees to focus more on high-value work. Other companies, see RPA as the next step before fully adopting intelligent automation technology such as machine learning and artificial intelligence. RPA is one of the fastest-growing sectors in the field of enterprise technology. In 2018 RPA software soared in value to $864 million, a growth of over 63%. In the course of this article, we clearly explain exactly what RPA really is and how it works. To help our understanding we will also explore the potential benefits and disadvantages of this technology. Finally, we will highlight some of the most powerful and exciting ways in which it is already transforming enterprises in a range of industries. Robotic Process Automation, or RPA for short, is a way of automating structured, repetitive, or rules-based tasks and processes. It has a number of different applications. Its tools can capture data, retrieve information, communicate with other digital systems and process transactions. Implementation can help to prevent human error, particularly when charged with completing long, repetitive tasks. It can also reduce labor costs. A report by Deloitte revealed that one large, commercial bank implemented RPA into 85 software bots. These were used to tackle 13 processes interacting with 1.5 million requests in a year.


A Novel Hybrid Framework for Hourly PM2.5 Concentration Forecasting Using CEEMDAN and Deep Temporal Convolutional Neural Network

arXiv.org Artificial Intelligence

For hourly PM2.5 concentration prediction, accurately capturing the data patterns of external factors that affect PM2.5 concentration changes, and constructing a forecasting model is one of efficient means to improve forecasting accuracy. In this study, a novel hybrid forecasting model based on complete ensemble empirical mode decomposition with adaptive noise (CEEMDAN) and deep temporal convolutional neural network (DeepTCN) is developed to predict PM2.5 concentration, by modelling the data patterns of historical pollutant concentrations data, meteorological data, and discrete time variables' data. Taking PM2.5 concentration of Beijing as the sample, experimental results showed that the forecasting accuracy of the proposed CEEMDAN-DeepTCN model is verified to be the highest when compared with the time series model, artificial neural network, and the popular deep learning models. The new model has improved the capability to model the PM2.5-related factor data patterns, and can be used as a promising tool for forecasting PM2.5 concentrations.


The Why, What and How of Artificial General Intelligence Chip Development

arXiv.org Artificial Intelligence

The AI chips increasingly focus on implementing neural computing at low power and cost. The intelligent sensing, automation, and edge computing applications have been the market drivers for AI chips. Increasingly, the generalisation, performance, robustness, and scalability of the AI chip solutions are compared with human-like intelligence abilities. Such a requirement to transit from application-specific to general intelligence AI chip must consider several factors. This paper provides an overview of this cross-disciplinary field of study, elaborating on the generalisation of intelligence as understood in building artificial general intelligence (AGI) systems. This work presents a listing of emerging AI chip technologies, classification of edge AI implementations, and the funnel design flow for AGI chip development. Finally, the design consideration required for building an AGI chip is listed along with the methods for testing and validating it.


A Number Sense as an Emergent Property of the Manipulating Brain

arXiv.org Artificial Intelligence

The ability to understand and manipulate numbers and quantities emerges during childhood, but the mechanism through which this ability is developed is still poorly understood. In particular, it is not known whether acquiring such a {\em number sense} is possible without supervision from a teacher. To explore this question, we propose a model in which spontaneous and undirected manipulation of small objects trains perception to predict the resulting scene changes. We find that, from this task, an image representation emerges that exhibits regularities that foreshadow numbers and quantity. These include distinct categories for zero and the first few natural numbers, a notion of order, and a signal that correlates with numerical quantity. As a result, our model acquires the ability to estimate the number of objects in the scene, as well as {\em subitization}, i.e. the ability to recognize at a glance the exact number of objects in small scenes. We conclude that important aspects of a facility with numbers and quantities may be learned without explicit teacher supervision.


An Approach to Intelligent Pneumonia Detection and Integration

arXiv.org Artificial Intelligence

Each year, over 2.5 million people, most of them in developed countries, die from pneumonia [1]. Since many studies have proved pneumonia is successfully treatable when timely and correctly diagnosed, many of diagnosis aids have been developed, with AI-based methods achieving high accuracies [2]. However, currently, the usage of AI in pneumonia detection is limited, in particular, due to challenges in generalizing a locally achieved result. In this report, we propose a roadmap for creating and integrating a system that attempts to solve this challenge. We also address various technical, legal, ethical, and logistical issues, with a blueprint of possible solutions.