Goto

Collaborating Authors

 Africa


Comparison of UAV and SAR performance for Crop type classification using machine learning algorithms: a case study of humid forest ecology experimental research site of West Africa

#artificialintelligence

Food insecurity is one of the major challenges facing African countries; therefore, timely and accurate information on agricultural production is essential to feed the growing population on the continent. A synergistic approach comprising a high-resolution multispectral UAV optical dataset and synthetic aperture radar (SAR) can help understand spectral features of target objects, especially with crop type identification. We conducted this work on the experimental plots using high spatial resolution multispectral UAV data (12 cm, re-sampled to 50 cm) in combination with the Sentinel 1C Synthetic Aperture Radar (SAR) dataset. Multiple combinations of the UAV datasets were analysed to assess the impact of canopy height model (CHM) on classification accuracy and to determine the optimum dataset (including spatial resolution) for the land cover classification. We also appraise the impact of variable spatial resolution on classification accuracy.


Eye-Tracker In The Car Keeps Drivers Awake And Alert

#artificialintelligence

A new generation of cars keeps an eye on you… to make sure you keep an eye on the road. A tiny camera on the dashboard monitors every blink of the driver's eyes to make sure they're not drowsy or distracted. It tracks the exact position and tilt of their face, the direction of gaze, eyelid activity, the rate and duration of every blink, how dilated their pupils are, how open their eyes are, whether their mouth is open, and more. Using AI and computer vision, it is constantly watching out for signs of cell phone usage, seatbelt-wearing and smoking, and checking that the driver is actually focused on the road. If they're not, it calls them out on it.


A Few Thousand Translations Go a Long Way! Leveraging Pre-trained Models for African News Translation

arXiv.org Artificial Intelligence

Recent advances in the pre-training of language models leverage large-scale datasets to create multilingual models. However, low-resource languages are mostly left out in these datasets. This is primarily because many widely spoken languages are not well represented on the web and therefore excluded from the large-scale crawls used to create datasets. Furthermore, downstream users of these models are restricted to the selection of languages originally chosen for pre-training. This work investigates how to optimally leverage existing pre-trained models to create low-resource translation systems for 16 African languages. We focus on two questions: 1) How can pre-trained models be used for languages not included in the initial pre-training? and 2) How can the resulting translation models effectively transfer to new domains? To answer these questions, we create a new African news corpus covering 16 languages, of which eight languages are not part of any existing evaluation dataset. We demonstrate that the most effective strategy for transferring both to additional languages and to additional domains is to fine-tune large pre-trained models on small quantities of high-quality translation data.


Are discrete units necessary for Spoken Language Modeling?

arXiv.org Artificial Intelligence

Recent work in spoken language modeling shows the possibility of learning a language unsupervisedly from raw audio without any text labels. The approach relies first on transforming the audio into a sequence of discrete units (or pseudo-text) and then training a language model directly on such pseudo-text. Is such a discrete bottleneck necessary, potentially introducing irreversible errors in the encoding of the speech signal, or could we learn a language model without discrete units at all? In this work, we study the role of discrete versus continuous representations in spoken language modeling. We show that discretization is indeed essential for good results in spoken language modeling. We show that discretization removes linguistically irrelevant information from the continuous features, helping to improve language modeling performances. On the basis of this study, we train a language model on the discrete units of the HuBERT features, reaching new state-of-the-art results in the lexical, syntactic and semantic metrics of the Zero Resource Speech Challenge 2021 (Track 1 - Speech Only).


Contributions \`a l'asservissement visuel et \`a l'imagerie en m\'edecine

arXiv.org Artificial Intelligence

This manuscript gives an overview of my research work carried out within the FEMTO-ST institute in Besan\c{c}on, more particularly in the Automatic and Micro-Mechatronic Systems (AS2M) department. It is above all the result of my (co)-supervision of interns, PhD students and postdocs. I would like to pay tribute to them, for their major contribution to scientific research, here and elsewhere.


Design Automation for Fast, Lightweight, and Effective Deep Learning Models: A Survey

arXiv.org Artificial Intelligence

Deep learning technologies have demonstrated remarkable effectiveness in a wide range of tasks, and deep learning holds the potential to advance a multitude of applications, including in edge computing, where deep models are deployed on edge devices to enable instant data processing and response. A key challenge is that while the application of deep models often incurs substantial memory and computational costs, edge devices typically offer only very limited storage and computational capabilities that may vary substantially across devices. These characteristics make it difficult to build deep learning solutions that unleash the potential of edge devices while complying with their constraints. A promising approach to addressing this challenge is to automate the design of effective deep learning models that are lightweight, require only a little storage, and incur only low computational overheads. This survey offers comprehensive coverage of studies of design automation techniques for deep learning models targeting edge computing. It offers an overview and comparison of key metrics that are used commonly to quantify the proficiency of models in terms of effectiveness, lightness, and computational costs. The survey then proceeds to cover three categories of the state-of-the-art of deep model design automation techniques: automated neural architecture search, automated model compression, and joint automated design and compression. Finally, the survey covers open issues and directions for future research.


FedSSO: A Federated Server-Side Second-Order Optimization Algorithm

arXiv.org Artificial Intelligence

In this work, we propose FedSSO, a server-side second-order optimization method for federated learning (FL). In contrast to previous works in this direction, we employ a server-side approximation for the Quasi-Newton method without requiring any training data from the clients. In this way, we not only shift the computation burden from clients to server, but also eliminate the additional communication for second-order updates between clients and server entirely. We provide theoretical guarantee for convergence of our novel method, and empirically demonstrate our fast convergence and communication savings in both convex and non-convex settings.


VL-InterpreT: An Interactive Visualization Tool for Interpreting Vision-Language Transformers

arXiv.org Artificial Intelligence

Breakthroughs in transformer-based models have revolutionized not only the NLP field, but also vision and multimodal systems. However, although visualization and interpretability tools have become available for NLP models, internal mechanisms of vision and multimodal transformers remain largely opaque. With the success of these transformers, it is increasingly critical to understand their inner workings, as unraveling these black-boxes will lead to more capable and trustworthy models. To contribute to this quest, we propose VL-InterpreT, which provides novel interactive visualizations for interpreting the attentions and hidden representations in multimodal transformers. VL-InterpreT is a task agnostic and integrated tool that (1) tracks a variety of statistics in attention heads throughout all layers for both vision and language components, (2) visualizes cross-modal and intra-modal attentions through easily readable heatmaps, and (3) plots the hidden representations of vision and language tokens as they pass through the transformer layers. In this paper, we demonstrate the functionalities of VL-InterpreT through the analysis of KD-VLP, an end-to-end pretraining vision-language multimodal transformer-based model, in the tasks of Visual Commonsense Reasoning (VCR) and WebQA, two visual question answering benchmarks. Furthermore, we also present a few interesting findings about multimodal transformer behaviors that were learned through our tool.


William MacAskill: 'There are 80 trillion people yet to come. They need us to start protecting them'

The Guardian

Although most cultures, particularly in the west, provide a great many commemorations of distant ancestors – statues, portraits, buildings – we are much less willing to consider our far-off descendants. We might invoke grandchildren, at a push great-grandchildren, but after that, it all becomes a bit vague and, well, unimaginable. And while we look with awe and fascination at the Egyptian pyramids, built 5,000 years ago, we seem incapable of thinking, or even contemplating, 5,000 years in the future. That lies in the realm of science fiction, which is tantamount to fantasy. But the chances are, barring a global catastrophe, humanity will still be very much around in 5,000 years, and going by the average existence of mammal species, should still be thriving in 500,000 years. If we play our cards right, we could even be here in 5m or 500m years, which means that there may be thousands or even millions times more human beings to come than have already existed.


Provably Tightest Linear Approximation for Robustness Verification of Sigmoid-like Neural Networks

arXiv.org Artificial Intelligence

The robustness of deep neural networks is crucial to modern AI-enabled systems and should be formally verified. Sigmoid-like neural networks have been adopted in a wide range of applications. Due to their non-linearity, Sigmoid-like activation functions are usually over-approximated for efficient verification, which inevitably introduces imprecision. Considerable efforts have been devoted to finding the so-called tighter approximations to obtain more precise verification results. However, existing tightness definitions are heuristic and lack theoretical foundations. We conduct a thorough empirical analysis of existing neuron-wise characterizations of tightness and reveal that they are superior only on specific neural networks. We then introduce the notion of network-wise tightness as a unified tightness definition and show that computing network-wise tightness is a complex non-convex optimization problem. We bypass the complexity from different perspectives via two efficient, provably tightest approximations. The results demonstrate the promising performance achievement of our approaches over state of the art: (i) achieving up to 251.28% improvement to certified lower robustness bounds; and (ii) exhibiting notably more precise verification results on convolutional networks.