Goto

Collaborating Authors

 Deep Learning


Learning Optimal Conformal Classifiers

arXiv.org Machine Learning

Modern deep learning based classifiers show very high accuracy on test data but this does not provide sufficient guarantees for safe deployment, especially in high-stake AI applications such as medical diagnosis. Usually, predictions are obtained without a reliable uncertainty estimate or a formal guarantee. Conformal prediction (CP) addresses these issues by using the classifier's probability estimates to predict confidence sets containing the true class with a user-specified probability. However, using CP as a separate processing step after training prevents the underlying model from adapting to the prediction of confidence sets. Thus, this paper explores strategies to differentiate through CP during training with the goal of training model with the conformal wrapper end-to-end. In our approach, conformal training (ConfTr), we specifically "simulate" conformalization on mini-batches during training. We show that CT outperforms state-of-the-art CP methods for classification by reducing the average confidence set size (inefficiency). Moreover, it allows to "shape" the confidence sets predicted at test time, which is difficult for standard CP. On experiments with several datasets, we show ConfTr can influence how inefficiency is distributed across classes, or guide the composition of confidence sets in terms of the included classes, while retaining the guarantees offered by CP.


Learning Prototype-oriented Set Representations for Meta-Learning

arXiv.org Machine Learning

Learning from set-structured data is a fundamental problem that has recently attracted increasing attention, where a series of summary networks are introduced to deal with the set input. In fact, many meta-learning problems can be treated as set-input tasks. Most existing summary networks aim to design different architectures for the input set in order to enforce permutation invariance. However, scant attention has been paid to the common cases where different sets in a meta-distribution are closely related and share certain statistical properties. Viewing each set as a distribution over a set of global prototypes, this paper provides a novel optimal transport (OT) based way to improve existing summary networks. To learn the distribution over the global prototypes, we minimize its OT distance to the set empirical distribution over data points, providing a natural unsupervised way to improve the summary network. Since our plug-and-play framework can be applied to many meta-learning problems, we further instantiate it to the cases of few-shot classification and implicit meta generative modeling. Extensive experiments demonstrate that our framework significantly improves the existing summary networks on learning more powerful summary statistics from sets and can be successfully integrated into metric-based few-shot classification and generative modeling applications, providing a promising tool for addressing set-input and meta-learning problems.


Natural Image Reconstruction from fMRI using Deep Learning: A Survey

arXiv.org Machine Learning

With the advent of brain imaging techniques and machine learning tools, much effort has been devoted to building computational models to capture the encoding of visual information in the human brain. One of the most challenging brain decoding tasks is the accurate reconstruction of the perceived natural images from brain activities measured by functional magnetic resonance imaging (fMRI). In this work, we survey the most recent deep learning methods for natural image reconstruction from fMRI. We examine these methods in terms of architectural design, benchmark datasets, and evaluation metrics and present a fair performance evaluation across standardized evaluation metrics. Finally, we discuss the strengths and limitations of existing studies and present potential future directions.


Astronomical source finding services for the CIRASA visual analytic platform

arXiv.org Machine Learning

Innovative developments in data processing, archiving, analysis, and visualization are nowadays unavoidable to deal with the data deluge expected in next-generation facilities for radio astronomy, such as the Square Kilometre Array (SKA) and its precursors. In this context, the integration of source extraction and analysis algorithms into data visualization tools could significantly improve and speed up the cataloguing process of large area surveys, boosting astronomer productivity and shortening publication time. To this aim, we are developing a visual analytic platform (CIRASA) for advanced source finding and classification, integrating state-of-the-art tools, such as the CAESAR source finder, the ViaLactea Visual Analytic (VLVA) and Knowledge Base (VLKB). In this work, we present the project objectives and the platform architecture, focusing on the implemented source finding services.


Decade Of Artificial Intelligence: A Summary

#artificialintelligence

The world has seen a boom in the field of Artificial Intelligence in the past few years. The major reasons contributing to this is the availability of data and computing power. A lot of research has happened in the field of AI in the last decade and society has witnessed many amazing use cases. In the last decade, AI went mainstream because of the availability of hardware, courses, platforms, big companies taking workshops, etc. What our AI community has achieved in the last decade has set a strong foundation for the future.


Deep Learning Enhances Cancer Diagnostic Tools

#artificialintelligence

Yi "Edwin" Sun, a Ph.D. candidate in electrical and computer engineering at the University of Illinois Urbana-Champaign and member of the Beckman Institute's Biophotonics Imaging Laboratory headed by Stephen Boppart, explored how deep learning methods can make polarization-sensitive optical coherence tomography, or PS-OCT, more cost-effective and better equipped to diagnose cancer in biological tissues. The paper, titled "Synthetic polarization-sensitive optical coherence tomography by deep learning," was published in npj Digital Medicine. OCT systems are common clinically and are used to generate high-resolution cross-sectional images of regions in the human body. Sun and his team developed a groundbreaking method of applying software to the OCT tool to provide polarization-sensitive capabilities -- without the cost and complexity that accompany hardware-based PS-OCT imaging systems. "We're trying to replace the hardware associated with PS-OCT," Sun said.


Rebuilding our next-gen ML Platform with the best of Spark and Tensorflow

#artificialintelligence

Rue Gilt Groupe is a fashion eCommerce company located in Boston, MA, that has 50M members and daily flash sales on millions of products. Our Data Science team is a tight-knit group of Data Scientists and Machine Learning Engineers who work full-stack on cloud-native architectures to deliver DS and ML services, heavily utilizing Apache Spark and AWS. This post focuses on some recent updates we incorporated into one of our stacks built for big data applications to add support for running the latest and greatest deep learning based algorithms and models. This architecture provides us with the flexibility to pick the right framework at any step of Machine Learning and unlock scalable deep learning pipelines with minimal MLOps code. At the same time, it also provides the flexibility to transition to any MLOps platform without a lot of future ML code changes.


AI and simulation tools to fight COVID-19

#artificialintelligence

In its on-going campaign to reveal the inner workings of the SARS-CoV-2 virus, the U.S. Department of Energy's (DOE) Argonne National Laboratory is leading efforts to couple artificial intelligence (AI) and cutting-edge simulation workflows to better understand biological observations and accelerate drug discovery. Argonne collaborated with academic and commercial research partners to achieve near real-time feedback between simulation and AI approaches to understand how two proteins in the SARS-CoV-2 viral genome, nsp10 and nsp16, interact to help the virus replicate and elude the host's immune system. The team achieved this milestone by coupling two distinct hardware platforms: Cerebras CS-1, a processor-packed silicon wafer deep learning accelerator; and ThetaGPU, an AI- and simulation-enabled extension of the Theta supercomputer, housed at the Argonne Leadership Computing Facility, a DOE Office of Science User Facility. To enable this capability, the team developed Stream-AI-MD, a novel application of the AI method called deep learning to drive adaptive molecular dynamics (MD) simulations in a streaming manner. Data from simulations is streamed from ThetaGPU onto the Cerebras CS-1 platform to simultaneously analyze how the two proteins interact.


AI Analysis of Bird Songs Helping Scientists Study Bird Populations and Movements - AI Trends

#artificialintelligence

A study of bird songs conducted in the Sierra Nevada mountain range in California generated a million hours of audio, which AI researchers are working to decode to gain insights into how birds responded to wildfires in the region, and to learn which measures helped the birds to rebound more quickly. Scientists can also use the soundscape to help track shifts in migration timing and population ranges, according to a recent account in Scientific American. More audio data is coming in from other research as well, with sound-based projects to count insects and study the effects of light and noise pollution on bird communities underway. "Audio data is a real treasure trove because it contains vast amounts of information," stated ecologist Connor Wood, a Cornell University postdoctoral researcher, who is leading the Sierra Nevada project. "We just need to think creatively about how to share and access that information."


AI Weekly: AI model training costs on the rise, highlighting need for new solutions

#artificialintelligence

This week, Microsoft and Nvidia announced that they trained what they claim is one of the largest and most capable AI language models to date: Megatron-Turing Natural Language Generation (MT-NLP). MT-NLP contains 530 billion parameters -- the parts of the model learned from historical data -- and achieves leading accuracy in a broad set of tasks, including reading comprehension and natural language inferences. But building it didn't come cheap. Experts peg the cost in the millions of dollars. Like other large AI systems, MT-NLP raises questions about the accessibility of cutting-edge research approaches in machine learning.