Africa
On Reinforcement Learning and Distribution Matching for Fine-Tuning Language Models with no Catastrophic Forgetting
Korbak, Tomasz, Elsahar, Hady, Kruszewski, Germán, Dymetman, Marc
The availability of large pre-trained models is changing the landscape of Machine Learning research and practice, moving from a training-from-scratch to a fine-tuning paradigm. While in some applications the goal is to "nudge" the pre-trained distribution towards preferred outputs, in others it is to steer it towards a different distribution over the sample space. Two main paradigms have emerged to tackle this challenge: Reward Maximization (RM) and, more recently, Distribution Matching (DM). RM applies standard Reinforcement Learning (RL) techniques, such as Policy Gradients, to gradually increase the reward signal. DM prescribes to first make explicit the target distribution that the model is fine-tuned to approximate. Here we explore the theoretical connections between the two paradigms, and show that methods such as KL-control developed for RM can also be construed as belonging to DM. We further observe that while DM differs from RM, it can suffer from similar training difficulties, such as high gradient variance. We leverage connections between the two paradigms to import the concept of baseline into DM methods. We empirically validate the benefits of adding a baseline on an array of controllable language generation tasks such as constraining topic, sentiment, and gender distributions in texts sampled from a language model. We observe superior performance in terms of constraint satisfaction, stability and sample efficiency.
Transfer learning and Local interpretable model agnostic based visual approach in Monkeypox Disease Detection and Classification: A Deep Learning insights
Ahsan, Md Manjurul, Abdullah, Tareque Abu, Ali, Md Shahin, Jahora, Fatematuj, Islam, Md Khairul, Alhashim, Amin G., Gupta, Kishor Datta
The recent development of Monkeypox disease among various nations poses a global pandemic threat when the world is still fighting Coronavirus Disease-2019 (COVID-19). At its dawn, the slow and steady transmission of Monkeypox disease among individuals needs to be addressed seriously. Over the years, Deep learning (DL) based disease prediction has demonstrated true potential by providing early, cheap, and affordable diagnosis facilities. Considering this opportunity, we have conducted two studies where we modified and tested six distinct deep learning models-VGG16, InceptionResNetV2, ResNet50, ResNet101, MobileNetV2, and VGG19-using transfer learning approaches. Our preliminary computational results show that the proposed modified InceptionResNetV2 and MobileNetV2 models perform best by achieving an accuracy ranging from 93% to 99%. Our findings are reinforced by recent academic work that demonstrates improved performance in constructing multiple disease diagnosis models using transfer learning approaches. Lastly, we further explain our model prediction using Local Interpretable Model-Agnostic Explanations (LIME), which play an essential role in identifying important features that characterize the onset of Monkeypox disease.
NFLAT: Non-Flat-Lattice Transformer for Chinese Named Entity Recognition
Wu, Shuang, Song, Xiaoning, Feng, Zhenhua, Wu, Xiao-Jun
Recently, Flat-LAttice Transformer (FLAT) has achieved great success in Chinese Named Entity Recognition (NER). FLAT performs lexical enhancement by constructing flat lattices, which mitigates the difficulties posed by blurred word boundaries and the lack of word semantics. In FLAT, the positions of starting and ending characters are used to connect a matching word. However, this method is likely to match more words when dealing with long texts, resulting in long input sequences. Therefore, it significantly increases the memory and computational costs of the self-attention module. To deal with this issue, we advocate a novel lexical enhancement method, InterFormer, that effectively reduces the amount of computational and memory costs by constructing non-flat lattices. Furthermore, with InterFormer as the backbone, we implement NFLAT for Chinese NER. NFLAT decouples lexicon fusion and context feature encoding. Compared with FLAT, it reduces unnecessary attention calculations in "word-character" and "word-word". This reduces the memory usage by about 50% and can use more extensive lexicons or higher batches for network training. The experimental results obtained on several well-known benchmarks demonstrate the superiority of the proposed method over the state-of-the-art hybrid (character-word) models.
Smart cities, smarter public health
Over the course of the last two years, we interviewed mayors, city officials, urban planners, academics, and citizens in cities around the world to identify the trends that are making urban living more sustainable, affordable, and human. One theme that emerged was cities' increasingly important role in ensuring the health and well-being of their residents.4 Cities currently represent just 3% of the world's territory but harbor 55% of the world's population. By 2050, it's estimated that 70% of the world's population will live in urban centers.5 At an economic level, cities generate around 80% of the global GDP,6 and are responsible for 80% of energy consumption and more than 70% of carbon emissions and global waste.7
Can robots impact our health? One study says so
A growing number of Americans are seeing their job security erode in the face of automation and it's undermining their health, according to a new study. The report, conducted by three Ball State University researchers with the school's Center for Business and Economic Research, shows that a 10 percentage point increase in automation risk increases average per-county costs related to medical expenses and lost productivity. "People who live and work in areas where automation is taking place are sickened by the thought of losing their jobs and having no way of providing for themselves or their families," said Michael Hicks, the center's director, who helped conduct the research. Costs associated with an increase in poor or fair health rise by $24 million to $174 million, costs related to increased physical distress rise by $6 million to $40 million and costs linked to mental distress increase by $7 million to $47 million. "This should give us pause about thinking through the benefits and costs of these technologies," Hicks said.
Hashish and pirates: How AI is cleaning up the high seas
On August 8th, 2021, Spanish police and customs agents intercepted the cargo ship NATALIA on suspicion of narcotics trafficking. The ship was en route from Lebanon via Iskenderun, Turkey, to Lagos, Nigeria, and hidden on board was nearly 20 tons of hashish worth $470 million. That may sound like the opening scene of an action flick, but it's the kind of occurrence that happens more frequently than you might expect on the high seas. Drug smuggling, illegal fishing, and piracy are constant threats. Following a number of recent piracy incidents in the Gulf of Aden, Iran, Russia, and China recently began naval and air drills seeking to counter maritime piracy.
Deep Learning-enabled Virtual Histological Staining of Biological Samples
Bai, Bijie, Yang, Xilin, Li, Yuzhu, Zhang, Yijie, Pillar, Nir, Ozcan, Aydogan
Histological staining is the gold standard for tissue examination in clinical pathology and life-science research, which visualizes the tissue and cellular structures using chromatic dyes or fluorescence labels to aid the microscopic assessment of tissue. However, the current histological staining workflow requires tedious sample preparation steps, specialized laboratory infrastructure, and trained histotechnologists, making it expensive, time-consuming, and not accessible in resource-limited settings. Deep learning techniques created new opportunities to revolutionize staining methods by digitally generating histological stains using trained neural networks, providing rapid, cost-effective, and accurate alternatives to standard chemical staining methods. These techniques, broadly referred to as virtual staining, were extensively explored by multiple research groups and demonstrated to be successful in generating various types of histological stains from label-free microscopic images of unstained samples; similar approaches were also used for transforming images of an already stained tissue sample into another type of stain, performing virtual stain-to-stain transformations. In this Review, we provide a comprehensive overview of the recent research advances in deep learning-enabled virtual histological staining techniques. The basic concepts and the typical workflow of virtual staining are introduced, followed by a discussion of representative works and their technical innovations. We also share our perspectives on the future of this emerging field, aiming to inspire readers from diverse scientific fields to further expand the scope of deep learning-enabled virtual histological staining techniques and their applications.
SPE: Symmetrical Prompt Enhancement for Fact Probing
Li, Yiyuan, Che, Tong, Wang, Yezhen, Jiang, Zhengbao, Xiong, Caiming, Chaturvedi, Snigdha
Pretrained language models (PLMs) have been shown to accumulate factual knowledge during pretrainingng (Petroni et al., 2019). Recent works probe PLMs for the extent of this knowledge through prompts either in discrete or continuous forms. However, these methods do not consider symmetry of the task: object prediction and subject prediction. In this work, we propose Symmetrical Prompt Enhancement (SPE), a continuous prompt-based method for factual probing in PLMs that leverages the symmetry of the task by constructing symmetrical prompts for subject and object prediction. Our results on a popular factual probing dataset, LAMA, show significant improvement of SPE over previous probing methods.
Methods for Recovering Conditional Independence Graphs: A Survey
Shrivastava, Harsh, Chajewska, Urszula
Conditional Independence (CI) graphs are a type of probabilistic graphical models that are primarily used to gain insights about feature relationships. Each edge represents the partial correlation between the connected features which gives information about their direct dependence. In this survey, we list out different methods and study the advances in techniques developed to recover CI graphs. We cover traditional optimization methods as well as recently developed deep learning architectures along with their recommended implementations. To facilitate wider adoption, we include preliminaries that consolidate associated operations, for example techniques to obtain covariance matrix for mixed datatypes. It is often beneficial to know which features are directly correlated to which other features. This can help us understand the input data better by giving a feature inter-dependence overview and also assist in taking system design decisions.
Elliptically-Contoured Tensor-variate Distributions with Application to Improved Image Learning
Llosa-Vite, Carlos, Maitra, Ranjan
Statistical analysis of tensor-valued data has largely used the tensor-variate normal (TVN) distribution that may be inadequate when data comes from distributions with heavier or lighter tails. We study a general family of elliptically contoured (EC) tensor-variate distributions and derive its characterizations, moments, marginal and conditional distributions, and the EC Wishart distribution. We describe procedures for maximum likelihood estimation from data that are (1) uncorrelated draws from an EC distribution, (2) from a scale mixture of the TVN distribution, and (3) from an underlying but unknown EC distribution, where we extend Tyler's robust estimator. A detailed simulation study highlights the benefits of choosing an EC distribution over the TVN for heavier-tailed data. We develop tensor-variate classification rules using discriminant analysis and EC errors and show that they better predict cats and dogs from images in the Animal Faces-HQ dataset than the TVN-based rules. A novel tensor-on-tensor regression and tensor-variate analysis of variance (TANOVA) framework under EC errors is also demonstrated to better characterize gender, age and ethnic origin than the usual TVN-based TANOVA in the celebrated Labeled Faces of the Wild dataset.