Africa
Establishing strong imputation performance of a denoising autoencoder in a wide range of missing data problems
Abiri, Najmeh, Linse, Björn, Edén, Patrik, Ohlsson, Mattias
Dealing with missing data in data analysis is inevitable. Although powerful imputation methods that address this problem exist, there is still much room for improvement. In this study, we examined single imputation based on deep autoencoders, motivated by the apparent success of deep learning to efficiently extract useful dataset features. We have developed a consistent framework for both training and imputation. Moreover, we benchmarked the results against state-of-the-art imputation methods on different data sizes and characteristics. The work was not limited to the one-type variable dataset; we also imputed missing data with multi-type variables, e.g., a combination of binary, categorical, and continuous attributes. To evaluate the imputation methods, we randomly corrupted the complete data, with varying degrees of corruption, and then compared the imputed and original values. In all experiments, the developed autoencoder obtained the smallest error for all ranges of initial data corruption.
Moroccan Artificial Intelligence Expert Joins UNESCO Ethics Commission
UNESCO has appointed Moroccan artificial intelligence expert, Mrs. Amal El Fallah Seghrouchni, to the World Commission on the Ethics of Scientific Knowledge and Technology (COMEST). The Moroccan researcher joins the commission for a four-year term, from 2020 to 2023. "It is an honor for me to serve ethics within this beautiful institution that is UNESCO," Seghrouchni shared on Twitter. The researcher holds a doctorate in artificial intelligence from the Pierre and Marie Curie University in Paris. She is professor at the School of Science and Engineering of Sorbonne University.
Semi-supervised acoustic and language model training for English-isiZulu code-switched speech recognition
Biswas, A., de Wet, F., van der Westhuizen, E., Niesler, T. R.
We present an analysis of semi-supervised acoustic and language model training for English-isiZulu code-switched ASR using soap opera speech. Approximately 11 hours of untranscribed multilingual speech was transcribed automatically using four bilingual code-switching transcription systems operating in English-isiZulu, English-isiXhosa, English-Setswana and English-Sesotho. These transcriptions were incorporated into the acoustic and language model training sets. Results showed that the TDNN-F acoustic models benefit from the additional semi-supervised data and that even better performance could be achieved by including additional CNN layers. Using these CNN-TDNN-F acoustic models, a first iteration of semi-supervised training achieved an absolute mixed-language WER reduction of 3.4%, and a further 2.2% after a second iteration. Although the languages in the untranscribed data were unknown, the best results were obtained when all automatically transcribed data was used for training and not just the utterances classified as English-isiZulu. Despite reducing perplexity, the semi-supervised language model was not able to improve the ASR performance.
GraphChallenge.org Sparse Deep Neural Network Performance
Kepner, Jeremy, Alford, Simon, Gadepally, Vijay, Jones, Michael, Milechin, Lauren, Reuther, Albert, Robinett, Ryan, Samsi, Sid
The MIT/IEEE/Amazon GraphChallenge.org encourages community approaches to developing new solutions for analyzing graphs and sparse data. Sparse AI analytics present unique scalability difficulties. The Sparse Deep Neural Network (DNN) Challenge draws upon prior challenges from machine learning, high performance computing, and visual analytics to create a challenge that is reflective of emerging sparse AI systems. The sparse DNN challenge is based on a mathematically well-defined DNN inference computation and can be implemented in any programming environment. In 2019 several sparse DNN challenge submissions were received from a wide range of authors and organizations. This paper presents a performance analysis of the best performers of these submissions. These submissions show that their state-of-the-art sparse DNN execution time, $T_{\rm DNN}$, is a strong function of the number of DNN operations performed, $N_{\rm op}$. The sparse DNN challenge provides a clear picture of current sparse DNN systems and underscores the need for new innovations to achieve high performance on very large sparse DNNs.
TAPAS: Weakly Supervised Table Parsing via Pre-training
Herzig, Jonathan, Nowak, Paweł Krzysztof, Müller, Thomas, Piccinno, Francesco, Eisenschlos, Julian Martin
Answering natural language questions over tables is usually seen as a semantic parsing task. To alleviate the collection cost of full logical forms, one popular approach focuses on weak supervision consisting of denotations instead of logical forms. However, training semantic parsers from weak supervision poses difficulties, and in addition, the generated logical forms are only used as an intermediate step prior to retrieving the denotation. In this paper, we present TAPAS, an approach to question answering over tables without generating logical forms. TAPAS trains from weak supervision, and predicts the denotation by selecting table cells and optionally applying a corresponding aggregation operator to such selection. TAPAS extends BERT's architecture to encode tables as input, initializes from an effective joint pre-training of text segments and tables crawled from Wikipedia, and is trained end-to-end. We experiment with three different semantic parsing datasets, and find that TAPAS outperforms or rivals semantic parsing models by improving state-of-the-art accuracy on SQA from 55.1 to 67.2 and performing on par with the state-of-the-art on WIKISQL and WIKITQ, but with a simpler model architecture. We additionally find that transfer learning, which is trivial in our setting, from WIKISQL to WIKITQ, yields 48.7 accuracy, 4.2 points above the state-of-the-art.
Show me your ID: Tunisia deploys 'robocop' to enforce COVID-19 lockdown
Tunisia deployed a police robot to patrol streets of the capital and enforce a lockdown imposed to contain coronavirus spread. Known as PGuard, the "robocop" which is remotely operated and is equipped with thermal imaging cameras is seen calling out to suspected violators in a video, "What are you doing? You don't know there's a lockdown?"
Men who are 'couch potatoes' are more likely to want to be muscular and hit the gym, study shows
Experts found that men from wealthy western countries like the UK are more motivated to workout than their Nicaraguan and Ugandan counterparts. However, in all three countries, men that watch more television -- and are therefore exposed more to images of idealised bodies -- wanted to be muscular more. Men who are'couch potatoes' -- those spending a lot of time watching TV -- are more likely to want to be muscular and hit the gym, a study has found Psychologist Tracey Thornborrow of the University of Lincoln and colleagues examined British men's obsession with getting a muscular physique -- along with related phenomena like relying on protein shakes, unhealthy dieting and steroid use. Comparing British men with those from Nicaragua and Uganda, the team assessed each man's body mass index, along with their feelings about peer pressure and their ideal appearance. Participants also ranked the perceived level of muscularity of their current body and their ideal body on the so-called'Male Adiposity and Muscularity Scale.' Designed by the Person Perception Lab at the University of Lincoln, the new scale makes use of two-dimensional images created from 3D software, providing a more realistic range of body types and sizes based on measurements of real people.
Analysis of the COVID-19 pandemic by SIR model and machine learning technics for forecasting
Ndiaye, Babacar Mbaye, Tendeng, Lena, Seck, Diaraf
This work is a trial in which we propose SIR model and machine learning tools to analyze the coronavirus pandemic in the real world. Based on the public data from \cite{datahub}, we estimate main key pandemic parameters and make predictions on the inflection point and possible ending time for the real world and specifically for Senegal. The coronavirus disease 2019, by World Health Organization, rapidly spread out in the whole China and then in the whole world. Under optimistic estimation, the pandemic in some countries will end soon, while for most part of countries in the world (US, Italy, etc.), the hit of anti-pandemic will be no later than the end of April.
Machine Learning the Phenomenology of COVID-19 From Early Infection Dynamics
We present a robust data-driven machine learning analysis of the COVID-19 pandemic from its early infection dynamics, specifically infection counts over time. The goal is to extract actionable public health insights. These insights include the infectious force, the rate of a mild infection becoming serious, estimates for asymtomatic infections and predictions of new infections over time. We focus on USA data starting from the first confirmed infection on January 20 2020. Our methods reveal significant asymptomatic (hidden) infection, a lag of about 10 days, and we quantitatively confirm that the infectious force is strong with about a 0.14% transition from mild to serious infection. Our methods are efficient, robust and general, being agnostic to the specific virus and applicable to different populations or cohorts.
Modeling Rare Interactions in Time Series Data Through Qualitative Change: Application to Outcome Prediction in Intensive Care Units
Ibrahim, Zina, Wu, Honghan, Dobson, Richard
Many areas of research are characterised by the deluge of large-scale highly-dimensional time-series data. However, using the data available for prediction and decision making is hampered by the current lag in our ability to uncover and quantify true interactions that explain the outcomes.We are interested in areas such as intensive care medicine, which are characterised by i) continuous monitoring of multivariate variables and non-uniform sampling of data streams, ii) the outcomes are generally governed by interactions between a small set of rare events, iii) these interactions are not necessarily definable by specific values (or value ranges) of a given group of variables, but rather, by the deviations of these values from the normal state recorded over time, iv) the need to explain the predictions made by the model. Here, while numerous data mining models have been formulated for outcome prediction, they are unable to explain their predictions. We present a model for uncovering interactions with the highest likelihood of generating the outcomes seen from highly-dimensional time series data. Interactions among variables are represented by a relational graph structure, which relies on qualitative abstractions to overcome non-uniform sampling and to capture the semantics of the interactions corresponding to the changes and deviations from normality of variables of interest over time. Using the assumption that similar templates of small interactions are responsible for the outcomes (as prevalent in the medical domains), we reformulate the discovery task to retrieve the most-likely templates from the data.