Goto

Collaborating Authors

 Africa


Leveraging Data Collection and Unsupervised Learning for Code-switched Tunisian Arabic Automatic Speech Recognition

arXiv.org Artificial Intelligence

Crafting an effective Automatic Speech Recognition (ASR) solution for dialects demands innovative approaches that not only address the data scarcity issue but also navigate the intricacies of linguistic diversity. In this paper, we address the aforementioned ASR challenge, focusing on the Tunisian dialect. First, textual and audio data is collected and in some cases annotated. Second, we explore self-supervision, semi-supervision and few-shot code-switching approaches to push the state-of-the-art on different Tunisian test sets; covering different acoustic, linguistic and prosodic conditions. Finally, and given the absence of conventional spelling, we produce a human evaluation of our transcripts to avoid the noise coming from spelling inadequacies in our testing references. Our models, allowing to transcribe audio samples in a linguistic mix involving Tunisian Arabic, English and French, and all the data used during training and testing are released for public use and further improvements.


How Much Temporal Long-Term Context is Needed for Action Segmentation?

arXiv.org Artificial Intelligence

Modeling long-term context in videos is crucial for many fine-grained tasks including temporal action segmentation. An interesting question that is still open is how much long-term temporal context is needed for optimal performance. While transformers can model the long-term context of a video, this becomes computationally prohibitive for long videos. Recent works on temporal action segmentation thus combine temporal convolutional networks with self-attentions that are computed only for a local temporal window. While these approaches show good results, their performance is limited by their inability to capture the full context of a video. In this work, we try to answer how much long-term temporal context is required for temporal action segmentation by introducing a transformer-based model that leverages sparse attention to capture the full context of a video. We compare our model with the current state of the art on three datasets for temporal action segmentation, namely 50Salads, Breakfast, and Assembly101. Our experiments show that modeling the full context of a video is necessary to obtain the best performance for temporal action segmentation.


FLARE: Fingerprinting Deep Reinforcement Learning Agents using Universal Adversarial Masks

arXiv.org Artificial Intelligence

We propose FLARE, the first fingerprinting mechanism to verify whether a suspected Deep Reinforcement Learning (DRL) policy is an illegitimate copy of another (victim) policy. We first show that it is possible to find non-transferable, universal adversarial masks, i.e., perturbations, to generate adversarial examples that can successfully transfer from a victim policy to its modified versions but not to independently trained policies. FLARE employs these masks as fingerprints to verify the true ownership of stolen DRL policies by measuring an action agreement value over states perturbed by such masks. Our empirical evaluations show that FLARE is effective (100% action agreement on stolen copies) and does not falsely accuse independent policies (no false positives). FLARE is also robust to model modification attacks and cannot be easily evaded by more informed adversaries without negatively impacting agent performance. We also show that not all universal adversarial masks are suitable candidates for fingerprints due to the inherent characteristics of DRL policies. The spatio-temporal dynamics of DRL problems and sequential decision-making process make characterizing the decision boundary of DRL policies more difficult, as well as searching for universal masks that capture the geometry of it.


How to estimate carbon footprint when training deep learning models? A guide and review

arXiv.org Artificial Intelligence

Machine learning and deep learning models have become essential in the recent fast development of artificial intelligence in many sectors of the society. It is now widely acknowledge that the development of these models has an environmental cost that has been analyzed in many studies. Several online and software tools have been developed to track energy consumption while training machine learning models. In this paper, we propose a comprehensive introduction and comparison of these tools for AI practitioners wishing to start estimating the environmental impact of their work. We review the specific vocabulary, the technical requirements for each tool. We compare the energy consumption estimated by each tool on two deep neural networks for image processing and on different types of servers. From these experiments, we provide some advice for better choosing the right tool and infrastructure.


GPU-based Private Information Retrieval for On-Device Machine Learning Inference

arXiv.org Artificial Intelligence

On-device machine learning (ML) inference can enable the use of private user data on user devices without revealing them to remote servers. However, a pure on-device solution to private ML inference is impractical for many applications that rely on embedding tables that are too large to be stored on-device. In particular, recommendation models typically use multiple embedding tables each on the order of 1-10 GBs of data, making them impractical to store on-device. To overcome this barrier, we propose the use of private information retrieval (PIR) to efficiently and privately retrieve embeddings from servers without sharing any private information. As off-the-shelf PIR algorithms are usually too computationally intensive to directly use for latency-sensitive inference tasks, we 1) propose novel GPU-based acceleration of PIR, and 2) co-design PIR with the downstream ML application to obtain further speedup. Our GPU acceleration strategy improves system throughput by more than $20 \times$ over an optimized CPU PIR implementation, and our PIR-ML co-design provides an over $5 \times$ additional throughput improvement at fixed model quality. Together, for various on-device ML applications such as recommendation and language modeling, our system on a single V100 GPU can serve up to $100,000$ queries per second -- a $>100 \times$ throughput improvement over a CPU-based baseline -- while maintaining model accuracy.


Deep Variational Free Energy Approach to Dense Hydrogen

arXiv.org Artificial Intelligence

Songshan Lake Materials Laboratory, Dongguan, Guangdong 523808, China (Dated: September 26, 2023) We developed a deep generative model-based variational free energy approach to the equations of state of dense hydrogen. We employ a normalizing flow network to model the proton Boltzmann distribution and a fermionic neural network to model the electron wave function at given proton positions. By jointly optimizing the two neural networks we reached a comparable variational free energy to the previous coupled electron-ion Monte Carlo calculation. The predicted equation of state of dense hydrogen under planetary conditions is denser than the findings of ab initio molecular dynamics calculation and empirical chemical model. Moreover, direct access to the entropy and free energy of dense hydrogen opens new opportunities in planetary modeling and high-pressure physics research. Hydrogen is the most abundant element in the visible universe.


Identification of Mixtures of Discrete Product Distributions in Near-Optimal Sample and Time Complexity

arXiv.org Machine Learning

We consider the problem of identifying, from statistics, a distribution of discrete random variables $X_1,\ldots,X_n$ that is a mixture of $k$ product distributions. The best previous sample complexity for $n \in O(k)$ was $(1/\zeta)^{O(k^2 \log k)}$ (under a mild separation assumption parameterized by $\zeta$). The best known lower bound was $\exp(\Omega(k))$. It is known that $n\geq 2k-1$ is necessary and sufficient for identification. We show, for any $n\geq 2k-1$, how to achieve sample complexity and run-time complexity $(1/\zeta)^{O(k)}$. We also extend the known lower bound of $e^{\Omega(k)}$ to match our upper bound across a broad range of $\zeta$. Our results are obtained by combining (a) a classic method for robust tensor decomposition, (b) a novel way of bounding the condition number of key matrices called Hadamard extensions, by studying their action only on flattened rank-1 tensors.


'Fox News Sunday' on September 24, 2023

FOX News

This is a rush transcript of'Fox News Sunday' on September 24, 2023. This copy may not be in its final form and may be updated. The chaos at the border grows by the day, as the pressure to take greater action builds yet again on the White House. We need people from the top. HEMMER (voice-over): A border city mayor and Democrat declaring a state of emergency as thousands upon thousands of migrants flow into the country. JOE BIDEN, PRESIDENT OF THE UNITED STATES: Republicans in Congress and my predecessor spent four years gutting the immigration system -- under my predecessor. They continue to undermine our border security today. HEMMER: We'll get reaction from border state Democrat, Texas Congressman Henry Cuellar. President Biden says he'll join the picket line in Michigan on Tuesday, just a day before Donald Trump will be there, too. Meanwhile, another presidential hopeful pushes back. TIM SCOTT (R-SC), PRESIDENTIAL CANDIDATE: We need a president who says we are not going to subsidize unions, period. HEMMER: We'll discuss with a man whose eyes are on the White House, South Carolina Senator Tim Scott. We'll ask Republican National Committee chairwoman Ronna McDaniel what voters can expect to see on stage Wednesday night. JAMES LANKFORD (R-OK): It's a symbol of respect for the country when you dress respectfully when you're doing this responsibility. JOHN FETTERMAN (D-PA): I think there are more important things we should be talking about rather if -- if I dressed like a slob. The number of illegals crossing our border hit another new record. We want to show you our FOX News drone camera from Eagle Pass, Texas. We've been watching remarkable images today of a human flood that shows no sign of receding. And today, a new survey shows how displeased Americans are with the president's border policies. In a moment, we'll speak with border state Democrat, Texas Congressman Henry Cuellar, on that. But, first, to Griff Jenkins who has been in Eagle Pass for what seems like several years now. Well, there's a humanitarian crisis playing out along our southern border in places like here in Eagle Pass, Texas, where migrants have traveled thousands of miles in hopes of reaching the U.S. in numbers far greater than what border officials are able to handle. Actions include sending active duty troops to the border, increasing deportations and granting temporary protective status to nearly half a million Venezuelans, making it easier for them to find work in cities like New York, where officials are struggling to find room for them. Meanwhile, Texas Governor Greg Abbott trying to deter the migrants from entering his state, with miles of dense razor wire, Humvees manning the riverbank and guardsmen in rafts attempting to turn them back.


Curiosity as a Self-Supervised Method to Improve Exploration in De novo Drug Design

arXiv.org Artificial Intelligence

In recent years, deep learning has demonstrated promising results in de novo drug design. However, the proposed techniques still lack an efficient exploration of the large chemical space. Most of these methods explore a small fragment of the chemical space of known drugs, if the desired molecules were not found, the process ends. In this work, we introduce a curiosity-driven method to force the model to navigate many parts of the chemical space, therefore, achieving higher desirability and diversity as well. At first, we train a recurrent neural network-based general molecular generator (G), then we fine-tune G to maximize curiosity and desirability. We define curiosity as the Tanimoto similarity between two generated molecules, a first molecule generated by G, and a second one generated by a copy of G (Gcopy). We only backpropagate the loss through G while keeping Gcopy unchanged. We benchmarked our approach against two desirable chemical properties related to drug-likeness and showed that the discovered chemical space can be significantly expanded, thus, discovering a higher number of desirable molecules with more diversity and potentially easier to synthesize. All Code and data used in this paper are available at https://github.com/amine179/Curiosity-RL-for-Drug-Design.


AspectCSE: Sentence Embeddings for Aspect-based Semantic Textual Similarity Using Contrastive Learning and Structured Knowledge

arXiv.org Artificial Intelligence

Generic sentence embeddings provide a coarse-grained approximation of semantic textual similarity but ignore specific aspects that make texts similar. Conversely, aspect-based sentence embeddings provide similarities between texts based on certain predefined aspects. Thus, similarity predictions of texts are more targeted to specific requirements and more easily explainable. In this paper, we present AspectCSE, an approach for aspect-based contrastive learning of sentence embeddings. Results indicate that AspectCSE achieves an average improvement of 3.97% on information retrieval tasks across multiple aspects compared to the previous best results. We also propose using Wikidata knowledge graph properties to train models of multi-aspect sentence embeddings in which multiple specific aspects are simultaneously considered during similarity predictions. We demonstrate that multi-aspect embeddings outperform single-aspect embeddings on aspect-specific information retrieval tasks. Finally, we examine the aspect-based sentence embedding space and demonstrate that embeddings of semantically similar aspect labels are often close, even without explicit similarity training between different aspect labels.