Africa
Climate change may already affect 85 percent of humanity: Report
Climate change could already be affecting 85 percent of the world's population, an analysis of tens of thousands of scientific studies has found. The analysis, released on Monday, was carried out by a team of researchers that used machine learning to comb through vast troves of research published between 1951 and 2018 and found some 100,000 papers that potentially documented evidence of climate change's effects on the Earth's systems. "We have overwhelming evidence that climate change is affecting all continents, all systems," study author Max Callaghan told the AFP news agency in an interview. He added there was a "huge amount of evidence" showing the ways in which these effects are being felt. The researchers taught a computer to identify climate-relevant studies, generating a list of papers on topics from disrupted butterfly migration to heat-related human deaths to forestry cover changes.
Representation of professions in entertainment media: Insights into frequency and sentiment trends through computational text analysis
Baruah, Sabyasachee, Somandepalli, Krishna, Narayanan, Shrikanth
Societal ideas and trends dictate media narratives and cinematic depictions which in turn influences people's beliefs and perceptions of the real world. Media portrayal of culture, education, government, religion, and family affect their function and evolution over time as people interpret and perceive these representations and incorporate them into their beliefs and actions. It is important to study media depictions of these social structures so that they do not propagate or reinforce negative stereotypes, or discriminate against any demographic section. In this work, we examine media representation of professions and provide computational insights into their incidence, and sentiment expressed, in entertainment media content. We create a searchable taxonomy of professional groups and titles to facilitate their retrieval from speaker-agnostic text passages like movie and television (TV) show subtitles. We leverage this taxonomy and relevant natural language processing (NLP) models to create a corpus of professional mentions in media content, spanning more than 136,000 IMDb titles over seven decades (1950-2017). We analyze the frequency and sentiment trends of different occupations, study the effect of media attributes like genre, country of production, and title type on these trends, and investigate if the incidence of professions in media subtitles correlate with their real-world employment statistics. We observe increased media mentions of STEM, arts, sports, and entertainment occupations in the analyzed subtitles, and a decreased frequency of manual labor jobs and military occupations. The sentiment expressed toward lawyers, police, and doctors is becoming negative over time, whereas astronauts, musicians, singers, and engineers are mentioned favorably. Professions that employ more people have increased media frequency, supporting our hypothesis that media acts as a mirror to society.
\beta-Intact-VAE: Identifying and Estimating Causal Effects under Limited Overlap
As an important problem in causal inference, we discuss the identification and estimation of treatment effects (TEs) under limited overlap; that is, when subjects with certain features belong to a single treatment group. We use a latent variable to model a prognostic score which is widely used in biostatistics and sufficient for TEs; i.e., we build a generative prognostic model. We prove that the latent variable recovers a prognostic score, and the model identifies individualized treatment effects. The model is then learned as \beta-Intact-VAE--a new type of variational autoencoder (VAE). We derive the TE error bounds that enable representations balanced for treatment groups conditioned on individualized features. The proposed method is compared with recent methods using (semi-)synthetic datasets.
Using UAVs for vehicle tracking and collision risk assessment at intersections
Zong, Shuya, Chen, Sikai, Alinizzi, Majed, Li, Yujie, Labi, Samuel
ABSTRACT Assessing collision risk is a critical challenge to effective traffic safety management. The deployment of unmanned aerial vehicles (UAVs) to address this issue has shown much promise, given their wide visual field and movement flexibility. This research demonstrates the application of UAVs and V2X connectivity to track the movement of road users and assess potential collisions at intersections. The study uses videos captured by UAVs. The proposed method combines deeplearning based tracking algorithms and time-to-collision tasks. The results not only provide beneficial information for vehicle's recognition of potential crashes and motion planning but also provided a valuable tool for urban road agencies and safety management engineers. INTRODUCTION It has been prognosticated that unmanned aerial vehicles (UAVs) will play a vital role in various application or context areas of transportation systems management. This is motivated by the success of UAVs in other domains including photography, photogrammetry, agriculture, terrain mapping, monitoring, disaster relief and rescue operations, and recreational purposes (1). Due to these applications, the emerging global market for drone-enabled services has been valued by the 2016 Middle East and North Africa Business Report at over $127B (2).
TCube: Domain-Agnostic Neural Time-series Narration
Sharma, Mandar, Brownstein, John S., Ramakrishnan, Naren
The task of generating rich and fluent narratives that aptly describe the characteristics, trends, and anomalies of time-series data is invaluable to the sciences (geology, meteorology, epidemiology) or finance (trades, stocks, or sales and inventory). The efforts for time-series narration hitherto are domain-specific and use predefined templates that offer consistency but lead to mechanical narratives. We present TCube (Time-series-to-text), a domain-agnostic neural framework for time-series narration, that couples the representation of essential time-series elements in the form of a dense knowledge graph and the translation of said knowledge graph into rich and fluent narratives through the transfer-learning capabilities of PLMs (Pre-trained Language Models). TCube's design primarily addresses the challenge that lies in building a neural framework in the complete paucity of annotated training data for time-series. The design incorporates knowledge graphs as an intermediary for the representation of essential time-series elements which can be linearized for textual translation. To the best of our knowledge, TCube is the first investigation of the use of neural strategies for time-series narration. Through extensive evaluations, we show that TCube can improve the lexical diversity of the generated narratives by up to 65.38% while still maintaining grammatical integrity. The practicality and deployability of TCube is further validated through an expert review (n=21) where 76.2% of participating experts wary of auto-generated narratives favored TCube as a deployable system for time-series narration due to its richer narratives. Our code-base, models, and datasets, with detailed instructions for reproducibility is publicly hosted at https://github.com/Mandar-Sharma/TCube.
Rome was built in 1776: A Case Study on Factual Correctness in Knowledge-Grounded Response Generation
Santhanam, Sashank, Hedayatnia, Behnam, Gella, Spandana, Padmakumar, Aishwarya, Kim, Seokhwan, Liu, Yang, Hakkani-Tur, Dilek
Recently neural response generation models have leveraged large pre-trained transformer models and knowledge snippets to generate relevant and informative responses. However, this does not guarantee that generated responses are factually correct. In this paper, we examine factual correctness in knowledge-grounded neural response generation models. We present a human annotation setup to identify three different response types: responses that are factually consistent with respect to the input knowledge, responses that contain hallucinated knowledge, and non-verifiable chitchat style responses. We use this setup to annotate responses generated using different stateof-the-art models, knowledge snippets, and decoding strategies. In addition, to facilitate the development of a factual consistency detector, we automatically create a new corpus called Conv-FEVER that is adapted from the Wizard of Wikipedia dataset and includes factually consistent and inconsistent responses. We demonstrate the benefit of our Conv-FEVER dataset by showing that the models trained on this data perform reasonably well to detect factually inconsistent responses with respect to the provided knowledge through evaluation on our human annotated data. We will release the Conv-FEVER dataset and the human annotated responses.
TEET! Tunisian Dataset for Toxic Speech Detection
Gharbi, Slim, Arfaoui, Heger, Haddad, Hatem, Kchaou, Mayssa
The complete freedom of expression in social media has its costs especially in spreading harmful and abusive content that may induce people to act accordingly. Therefore, the need of detecting automatically such a content becomes an urgent task that will help and enhance the efficiency in limiting this toxic spread. Compared to other Arabic dialects which are mostly based on MSA, the Tunisian dialect is a combination of many other languages like MSA, Tamazight, Italian and French. Because of its rich language, dealing with NLP problems can be challenging due to the lack of large annotated datasets. In this paper we are introducing a new annotated dataset composed of approximately 10k of comments. We provide an in-depth exploration of its vocabulary through feature engineering approaches as well as the results of the classification performance of machine learning classifiers like NB and SVM and deep learning models such as ARBERT, MARBERT and XLM-R.
Recurrent Model-Free RL is a Strong Baseline for Many POMDPs
Ni, Tianwei, Eysenbach, Benjamin, Salakhutdinov, Ruslan
Many problems in RL, such as meta RL, robust RL, and generalization in RL, can be cast as POMDPs. In theory, simply augmenting model-free RL with memory, such as recurrent neural networks, provides a general approach to solving all types of POMDPs. However, prior work has found that such recurrent model-free RL methods tend to perform worse than more specialized algorithms that are designed for specific types of POMDPs. This paper revisits this claim. We find that careful architecture and hyperparameter decisions yield a recurrent model-free implementation that performs on par with (and occasionally substantially better than) more sophisticated recent techniques in their respective domains. We also release a simple and efficient implementation of recurrent model-free RL for future work to use as a baseline for POMDPs. Code is available at https://github.com/twni2016/pomdp-baselines
Robust and Scalable SDE Learning: A Functional Perspective
Cameron, Scott, Cameron, Tyron, Pretorius, Arnu, Roberts, Stephen
Stochastic differential equations provide a rich class of flexible generative models, capable of describing a wide range of spatio-temporal processes. A host of recent work looks to learn data-representing SDEs, using neural networks and other flexible function approximators. Despite these advances, learning remains computationally expensive due to the sequential nature of SDE integrators. In this work, we propose an importance-sampling estimator for probabilities of observations of SDEs for the purposes of learning. Crucially, the approach we suggest does not rely on such integrators. The proposed method produces lower-variance gradient estimates compared to algorithms based on SDE integrators and has the added advantage of being embarrassingly parallelizable. Stochastic differential equations (SDEs) are a natural extension to ordinary differential equations which allows modelling of noisy and uncertain driving forces. These models are particularly appealing due to their flexibility in expressing highly complex relationships with simple equations, while retaining a high degree of interpretability. Much work has been done over the last century focussing on understanding and modelling with SDEs, particularly in dynamical systems and quantitative finance (Pavliotis, 2014; Malliavin & Thalmaier, 2006).
Drones Autonomously Attacked Humans for the First Time in March
The world's first recorded case of an autonomous drone attacking humans took place in March 2020, according to a United Nations (UN) security report detailing the ongoing Second Libyan Civil War. Libyan forces used the Turkish-made drones to "hunt down" and jam retreating enemy forces, preventing them from using their own drones. The field report (via New Scientist) describes how the Haftar Affiliated Forces (HAF), loyal to Libyan Field Marshal Khalifa Haftar, came under attack by drones from the rival Government of National Accord (GNA) forces. After a successful drive against HAF forces, the GNA launched drone attacks to press its advantage. The report says Turkey supplied the drones to Libyan forces, which is a violation of a UN arms embargo slapped on combatants in the conflict.