Goto

Collaborating Authors

 Africa


Robust Training of Neural Networks using Scale Invariant Architectures

arXiv.org Machine Learning

In contrast to SGD, adaptive gradient methods like Adam allow robust training of modern deep networks, especially large language models. However, the use of adaptivity not only comes at the cost of extra memory but also raises the fundamental question: can non-adaptive methods like SGD enjoy similar benefits? In this paper, we provide an affirmative answer to this question by proposing to achieve both robust and memory-efficient training via the following general recipe: (1) modify the architecture and make it scale invariant, i.e. the scale of parameter doesn't affect the output of the network, (2) train with SGD and weight decay, and optionally (3) clip the global gradient norm proportional to weight norm multiplied by $\sqrt{\tfrac{2\lambda}{\eta}}$, where $\eta$ is learning rate and $\lambda$ is weight decay. We show that this general approach is robust to rescaling of parameter and loss by proving that its convergence only depends logarithmically on the scale of initialization and loss, whereas the standard SGD might not even converge for many initializations. Following our recipe, we design a scale invariant version of BERT, called SIBERT, which when trained simply by vanilla SGD achieves performance comparable to BERT trained by adaptive methods like Adam on downstream tasks.


AI Research Associate for Early-Stage Scientific Discovery

arXiv.org Artificial Intelligence

Artificial intelligence (AI) has been increasingly applied in scientific activities for decades; however, it is still far from an insightful and trustworthy collaborator in the scientific process. Most existing AI methods are either too simplistic to be useful in real problems faced by scientists or too domain-specialized (even dogmatized), stifling transformative discoveries or paradigm shifts. We present an AI research associate for early-stage scientific discovery based on (a) a novel minimally-biased ontology for physics-based modeling that is context-aware, interpretable, and generalizable across classical and relativistic physics; (b) automatic search for viable and parsimonious hypotheses, represented at a high-level (via domain-agnostic constructs) with built-in invariants, e.g., postulated forms of conservation principles implied by a presupposed spacetime topology; and (c) automatic compilation of the enumerated hypotheses to domain-specific, interpretable, and trainable/testable tensor-based computation graphs to learn phenomenological relations, e.g., constitutive or material laws, from sparse (and possibly noisy) data sets.


Language Models Explain Word Reading Times Better Than Empirical Predictability

arXiv.org Artificial Intelligence

Though there is a strong consensus that word length and frequency are the most important single-word features determining visual-orthographic access to the mental lexicon, there is less agreement as how to best capture syntactic and semantic factors. The traditional approach in cognitive reading research assumes that word predictability from sentence context is best captured by cloze completion probability (CCP) derived from human performance data. We review recent research suggesting that probabilistic language models provide deeper explanations for syntactic and semantic effects than CCP. Then we compare CCP with (1) Symbolic n-gram models consolidate syntactic and semantic short-range relations by computing the probability of a word to occur, given two preceding words. (2) Topic models rely on subsymbolic representations to capture long-range semantic similarity by word co-occurrence counts in documents. (3) In recurrent neural networks (RNNs), the subsymbolic units are trained to predict the next word, given all preceding words in the sentences. To examine lexical retrieval, these models were used to predict single fixation durations and gaze durations to capture rapidly successful and standard lexical access, and total viewing time to capture late semantic integration. The linear item-level analyses showed greater correlations of all language models with all eye-movement measures than CCP. Then we examined non-linear relations between the different types of predictability and the reading times using generalized additive models. N-gram and RNN probabilities of the present word more consistently predicted reading performance compared with topic models or CCP.


An ASP approach for reasoning on neural networks under a finitely many-valued semantics for weighted conditional knowledge bases

arXiv.org Artificial Intelligence

Weighted knowledge bases for description logics with typicality have been recently considered under a "concept-wise" multipreference semantics (in both the two-valued and fuzzy case), as the basis of a logical semantics of MultiLayer Perceptrons (MLPs). In this paper we consider weighted conditional ALC knowledge bases with typicality in the finitely many-valued case, through three different semantic constructions, based on coherent, faithful and phi-coherent interpretations. For the boolean fragment LC of ALC we exploit ASP and "asprin" for reasoning with the concept-wise multipreference entailment under a phi-coherent semantics, suitable to characterize the stationary states of MLPs. As a proof of concept, we experiment the proposed approach for checking properties of trained MLPs.


FedSpace: An Efficient Federated Learning Framework at Satellites and Ground Stations

arXiv.org Machine Learning

Large-scale deployments of low Earth orbit (LEO) satellites collect massive amount of Earth imageries and sensor data, which can empower machine learning (ML) to address global challenges such as real-time disaster navigation and mitigation. However, it is often infeasible to download all the high-resolution images and train these ML models on the ground because of limited downlink bandwidth, sparse connectivity, and regularization constraints on the imagery resolution. To address these challenges, we leverage Federated Learning (FL), where ground stations and satellites collaboratively train a global ML model without sharing the captured images on the satellites. We show fundamental challenges in applying existing FL algorithms among satellites and ground stations, and we formulate an optimization problem which captures a unique trade-off between staleness and idleness. We propose a novel FL framework, named FedSpace, which dynamically schedules model aggregation based on the deterministic and time-varying connectivity according to satellite orbits. Extensive numerical evaluations based on real-world satellite images and satellite networks show that FedSpace reduces the training time by 1.7 days (38.6%) over the state-of-the-art FL algorithms.


Questions for Flat-Minima Optimization of Modern Neural Networks

arXiv.org Machine Learning

For training neural networks, flat-minima optimizers that seek to find parameters in neighborhoods having uniformly low loss (flat minima) have been shown to improve upon stochastic and adaptive gradient-based methods. Two methods for finding flat minima stand out: 1. Averaging methods (i.e., Stochastic Weight Averaging, SWA), and 2. Minimax methods (i.e., Sharpness Aware Minimization, SAM). However, despite similar motivations, there has been limited investigation into their properties and no comprehensive comparison between them. In this work, we investigate the loss surfaces from a systematic benchmarking of these approaches across computer vision, natural language processing, and graph learning tasks. The results lead to a simple hypothesis: since both approaches find different flat solutions, combining them should improve generalization even further. We verify this improves over either flat-minima approach in 39 out of 42 cases. When it does not, we investigate potential reasons. We hope our results across image, graph, and text data will help researchers to improve deep learning optimizers, and practitioners to pinpoint the optimizer for the problem at hand.


'Nothing to do, nowhere to go': What happens when elephants live alone

National Geographic

On a raw December day, as Christmas music blares over loudspeakers, an African elephant named Asha walks in tight circles in an enclosure at Natural Bridge Zoo, a roadside attraction in Virginia. Her living quarters consist of a barn and three outdoor yards--a fenced patch of grass about 90 by 40 feet, a dirt patch with a few logs scattered about, and a yard where she gives rides to children for $15 and her massive feet have worn a ring into the grass. Her space is barren--no shrubs, trees, or watering holes. Elephants, like humans, are social animals. In the wild, females typically live in herds of eight or more, yet Asha, who's nearly 40 years old, has been confined mostly alone for more than 30 years.


InstaDeep raises $100 million for decision support AI - Actu IA

#artificialintelligence

InstaDeep, one of the leaders in the design of decision-making Artificial Intelligence systems, announced on January 25 that it had raised $100 million (€88 million). The company closed a Series B round led by DeepTech investment firm Alpha Intelligence Capital and supported by CDIB. BioNTech, Chimera Abu Dhabi, Deutsche Bahn Digital Ventures, Google, G42 and Synergie participated in this latest round. Founded in 2014 by Karim Beguir and Zohra Slim, InstaDeep is a leader in decision AI systems, it has been named two years in a row to the CB Insights AI 100 ranking of the world's 100 most promising private artificial intelligence companies. The company develops patented AI products such as its DeepChainTM protein design platform and collaborates with leading companies such as Google DeepMind, Nvidia and Intel.


Jose Almeida on LinkedIn: #AI #africa #ai

#artificialintelligence

Organizations are looking deeper into data to gain a competitive advantage, implementing machine learning and artificial intelligence to achieve new business objectives and to move ahead of competitors in the industry. The adoption of AI and machine learning is critically impaired by the necessity of high-quality data. Changes must be made organization wide to identify and reduce pouches of bad data and create mechanisms that allow the organization to adapt quickly to the data needs and embrace the full potential of these technologies. From starting to being able to deliver a successful AI strategy goes a distance. The capability to build a secure, centralized, and scalable data repository, being able to combine large volumes of disparate data from multiples data sources, is the first challenge to overcome.


Artificial intelligence can be used to tackle COVID-19 inequities

#artificialintelligence

TORONTO, Jan. 31, 2022 – Artificial Intelligence (AI) can help tackle inequities that contribute to a higher risk of the most vulnerable contracting and dying of COVID-19, but York University researchers say the right data is crucial for that to happen. Vulnerable people are often more exposed to COVID-19 through their work, such as meat packing plants, and their living conditions which are often crowed, and they face more economic barriers, such having to rely on public transportation. York University Assistant Professor Jude Kong, Associate Professor Ali Asgary, and Distinguished Research Professor Jianhong Wu, can discuss how AI can play a role in eliminating inequities, especially during crises such as the current pandemic, ahead of upcoming webinar – Discovering COVID-19 Inequities and Systemic Vulnerabilities the Role of Artificial Intelligent Policy Implications. The webinar is part of the Transformative Disaster Risk Governance Webinar Series. "There is a need to use artificial intelligence to collect data disaggregated by race, gender, sexuality, class, geographic location and Indigeneity to better understand how COVID-19 is disproportionately affecting vulnerable people, whether here in Canada or in Africa, where many countries have difficulty obtaining vaccines. This kind of data could not only help with today's pandemic, but prepare for future crises by ensuring effective allocation of resources," says Kong, Faculty of Science, and director of the Africa-Canada Artificial Intelligence and Data Innovation Consortium.