Africa
How AI super-powers finance systems - Sage Advice South Africa
Digital technologies are transforming every aspect of our professional and personal lives. While certainly disruptive, and often the brunt of doomsday futurists, innovations like artificial intelligence (AI) present enormous possibilities to make things work better. They also add some exciting sparkle to traditionally dour (but very important) professions such as financial management. Shiny new image aside, AI is set to transform the financial management sector into a body of strategic, informed, and responsive consultants. There is no doubt that the COVID-19 pandemic has accelerated digital transformation, encouraging many finance professionals to harness the benefits of AI to meet their clients' changing expectations. The well-timed Sage CFO 3.0 research surveyed South African CFOs and other senior financial decision-makers for their views on the role of digital technologies, such as AI, in navigating uncertainty and preparing for the new working world.
These Algorithms Are Hunting for an EV Battery Mother Lode
"These things are hard to tip over," geologist Wilson Bonner assures me as the four-wheeled all-terrain vehicle he's piloting tilts suddenly sideways, pitching me toward the churned up mud beneath our wheels. We're grinding up the side of a thickly forested hill in rural Ontario, Canada, on a chilly fall day, heading toward a spot that Bonner's employer, startup KoBold Metals, says represents the marriage of cutting-edge artificial intelligence with one of humanity's oldest industries. We do indeed complete the half-hour trek relatively unmuddied, finally breaking through a ring of broken trees and mangled brush into a swath of bulldozed mud. A black pipe about as wide around as my arm juts out of the ground--the top end of a hole nearly a kilometer deep that was punched into the ground by a truck-sized drilling rig that sits idly nearby. It's not much to look at, but this hole might mark a step into the future of mining, an industry crucial for the world's transition to renewable energy.
Human Mobility Modeling During the COVID-19 Pandemic via Deep Graph Diffusion Infomax
Liu, Yang, Rong, Yu, Guo, Zhuoning, Chen, Nuo, Xu, Tingyang, Tsung, Fugee, Li, Jia
Non-Pharmaceutical Interventions (NPIs), such as social gathering restrictions, have shown effectiveness to slow the transmission of COVID-19 by reducing the contact of people. To support policy-makers, multiple studies have first modeled human mobility via macro indicators (e.g., average daily travel distance) and then studied the effectiveness of NPIs. In this work, we focus on mobility modeling and, from a micro perspective, aim to predict locations that will be visited by COVID-19 cases. Since NPIs generally cause economic and societal loss, such a micro perspective prediction benefits governments when they design and evaluate them. However, in real-world situations, strict privacy data protection regulations result in severe data sparsity problems (i.e., limited case and location information). To address these challenges, we formulate the micro perspective mobility modeling into computing the relevance score between a diffusion and a location, conditional on a geometric graph. we propose a model named Deep Graph Diffusion Infomax (DGDI), which jointly models variables including a geometric graph, a set of diffusions and a set of locations.To facilitate the research of COVID-19 prediction, we present two benchmarks that contain geometric graphs and location histories of COVID-19 cases. Extensive experiments on the two benchmarks show that DGDI significantly outperforms other competing methods.
On Computing Probabilistic Abductive Explanations
Izza, Yacine, Huang, Xuanxiang, Ignatiev, Alexey, Narodytska, Nina, Cooper, Martin C., Marques-Silva, Joao
The most widely studied explainable AI (XAI) approaches are unsound. This is the case with well-known model-agnostic explanation approaches, and it is also the case with approaches based on saliency maps. One solution is to consider intrinsic interpretability, which does not exhibit the drawback of unsoundness. Unfortunately, intrinsic interpretability can display unwieldy explanation redundancy. Formal explainability represents the alternative to these non-rigorous approaches, with one example being PI-explanations. Unfortunately, PI-explanations also exhibit important drawbacks, the most visible of which is arguably their size. Recently, it has been observed that the (absolute) rigor of PI-explanations can be traded off for a smaller explanation size, by computing the so-called relevant sets. Given some positive {\delta}, a set S of features is {\delta}-relevant if, when the features in S are fixed, the probability of getting the target class exceeds {\delta}. However, even for very simple classifiers, the complexity of computing relevant sets of features is prohibitive, with the decision problem being NPPP-complete for circuit-based classifiers. In contrast with earlier negative results, this paper investigates practical approaches for computing relevant sets for a number of widely used classifiers that include Decision Trees (DTs), Naive Bayes Classifiers (NBCs), and several families of classifiers obtained from propositional languages. Moreover, the paper shows that, in practice, and for these families of classifiers, relevant sets are easy to compute. Furthermore, the experiments confirm that succinct sets of relevant features can be obtained for the families of classifiers considered.
Tensor-based Sequential Learning via Hankel Matrix Representation for Next Item Recommendations
Frolov, Evgeny, Oseledets, Ivan
Self-attentive transformer models have recently been shown to solve the next item recommendation task very efficiently. The learned attention weights capture sequential dynamics in user behavior and generalize well. Motivated by the special structure of learned parameter space, we question if it is possible to mimic it with an alternative and more lightweight approach. We develop a new tensor factorization-based model that ingrains the structural knowledge about sequential data within the learning process. We demonstrate how certain properties of a self-attention network can be reproduced with our approach based on special Hankel matrix representation. The resulting model has a shallow linear architecture and compares competitively to its neural counterpart.
Real-World Compositional Generalization with Disentangled Sequence-to-Sequence Learning
Compositional generalization is a basic mechanism in human language learning, which current neural networks struggle with. A recently proposed Disentangled sequence-to-sequence model (Dangle) shows promising generalization capability by learning specialized encodings for each decoding step. We introduce two key modifications to this model which encourage more disentangled representations and improve its compute and memory efficiency, allowing us to tackle compositional generalization in a more realistic setting. Specifically, instead of adaptively re-encoding source keys and values at each time step, we disentangle their representations and only re-encode keys periodically, at some interval. Our new architecture leads to better generalization performance across existing tasks and datasets, and a new machine translation benchmark which we create by detecting naturally occurring compositional patterns in relation to a training set. We show this methodology better emulates real-world requirements than artificial challenges.
A Learning and Control Perspective for Microfinance
Kurniawan, Christian, Deng, Xiyu, Chakraborty, Adhiraj, Gueye, Assane, Chen, Niangjun, Nakahira, Yorie
Microfinance, despite its significant potential for poverty reduction, is facing sustainability hardships due to high default rates. Although many methods in regular finance can estimate credit scores and default probabilities, these methods are not directly applicable to microfinance due to the following unique characteristics: a) under-explored (developing) areas such as rural Africa do not have sufficient prior loan data for microfinance institutions (MFIs) to establish a credit scoring system; b) microfinance applicants may have difficulty providing sufficient information for MFIs to accurately predict default probabilities; and c) many MFIs use group liability (instead of collateral) to secure repayment. Here, we present a novel control-theoretic model of microfinance that accounts for these characteristics. We construct an algorithm to learn microfinance decision policies that achieve financial inclusion, fairness, social welfare, and sustainability. We characterize the convergence conditions to Pareto-optimum and the convergence speeds. We demonstrate, in numerous real and synthetic datasets, that the proposed method accounts for the complexities induced by group liability to produce robust decisions before sufficient loans are given to establish credit scoring systems and for applicants whose default probability cannot be accurately estimated due to missing information. To the best of our knowledge, this paper is the first to connect microfinance and control theory. We envision that the connection will enable safe learning and control techniques to help modernize microfinance and alleviate poverty.
Quantum Phase Recognition using Quantum Tensor Networks
Sahoo, Shweta, Azad, Utkarsh, Singh, Harjinder
Machine learning (ML) has recently facilitated many advances in solving problems related to many-body physical systems. Given the intrinsic quantum nature of these problems, it is natural to speculate that quantum-enhanced machine learning will enable us to unveil even greater details than we currently have. With this motivation, this paper examines a quantum machine learning approach based on shallow variational ansatz inspired by tensor networks for supervised learning tasks. In particular, we first look at the standard image classification tasks using the Fashion-MNIST dataset and study the effect of repeating tensor network layers on ansatz's expressibility and performance. Finally, we use this strategy to tackle the problem of quantum phase recognition for the transverse-field Ising and Heisenberg spin models in one and two dimensions, where we were able to reach $\geq 98\%$ test-set accuracies with both multi-scale entanglement renormalization ansatz (MERA) and tree tensor network (TTN) inspired parametrized quantum circuits.
DziriBERT: a Pre-trained Language Model for the Algerian Dialect
Abdaoui, Amine, Berrimi, Mohamed, Oussalah, Mourad, Moussaoui, Abdelouahab
Pre-trained transformers are now the de facto models in Natural Language Processing given their state-of-the-art results in many tasks and languages. However, most of the current models have been trained on languages for which large text resources are already available (such as English, French, Arabic, etc.). Therefore, there are still a number of low-resource languages that need more attention from the community. In this paper, we study the Algerian dialect which has several specificities that make the use of Arabic or multilingual models inappropriate. To address this issue, we collected more than one million Algerian tweets, and pre-trained the first Algerian language model: DziriB-ERT. When compared with existing models, DziriBERT achieves better results, especially when dealing with the Roman script. The obtained results show that pre-training a dedicated model on a small dataset (150 MB) can outperform existing models that have been trained on much more data (hundreds of GB). Finally, our model is publicly available to the community.
A Survey of Graph Neural Networks for Social Recommender Systems
Sharma, Kartik, Lee, Yeon-Chang, Nambi, Sivagami, Salian, Aditya, Shah, Shlok, Kim, Sang-Wook, Kumar, Srijan
Exploiting social relations in recommendation works well because of the effects of social homophily [61] and social influence [60]: (1) social homophily indicates that a user tends to connect herself to other users with similar attributes and preferences, and (2) social influence indicates that users with direct or indirect relations tend to influence each other to make themselves become more similar. Accordingly, SocialRS can effectively mitigate the data sparsity problem by exploiting social neighbors to capture the preferences of a sparsely interacting user. Literature has shown that SocialRS can be applied successfully in various recommendation domains (e.g., product [101, 103], music [116-118], location [39, 72, 100], and image [86, 99, 102]), thereby improving user satisfaction. Furthermore, techniques and insights explored from SocialRS can also be exploited in real-world applications other than recommendations. For instance, García-Sánchez et al. [20] leveraged SocialRS to design a decision-making system for marketing (e.g., advertisement), while Gasparetti et al. [21] analyzed SocialRS in terms of community detection. Motivated by such wide applicability, there has been an increasing interest in research on developing accurate 40 SocialRS models. In the early days, research focused on matrix factorization (MF) techniques [28, 54-20 57, 84, 112].