Europe
Band Selection from Hyperspectral Images Using Attention-based Convolutional Neural Networks
Lorenzo, Pablo Ribalta, Tulczyjew, Lukasz, Marcinkiewicz, Michal, Nalepa, Jakub
Abstract--This paper introduces new attention-based convolutional neural networks for selecting bands from hyperspectral images. The proposed approach reuses convolutional activations at different depths, identifying the most informative regions of the spectrum with the help of gating mechanisms. Our attention techniques are modular and easy to implement, and they can be seamlessly trained end-to-end using gradient descent. Our rigorous experiments showed that deep models equipped with the attention mechanism deliver high-quality classification, and repeatedly identify significant bands in the training data, permitting the creation of refined and extremely compact sets that retain the most meaningful features. Hyperspectral data's high dimensionality is an important challenge towards its accurate segmentation, efficient analysis, transfer and storage.
A Deep Learning Mechanism for Efficient Information Dissemination in Vehicular Floating Content
Manzo, Gaetano, Montenegro, Juan Sebastian Otรกlora, Rizzo, Gianluca
Abstract--Handling the tremendous amount of network data, produced by the explosive growth of mobile traffic volume, is becoming of main priority to achieve desired performance targets efficiently. Opportunistic communication such as Floating Content (FC), can be used to offload part of the cellular traffic volume to vehicular-to-vehicular communication (V2V), leaving to the infrastructure the task of coordinating the communication. Existing FC dimensioning approaches have limitations, mainly due to unrealistic assumptions and on a coarse partitioning of users, which results in over-dimensioning. Shaping the opportunistic communication area is a crucial task to achieve desired application performance efficiently. In this work, we propose a solution for this open challenge. In particular, the broadcasting areas called Anchor Zone (AZ), are selected via a deep learning approach to minimize communication resources achieving desired message availability. No assumption required to fit the classifier in both synthetic and real mobility. A numerical study is made to validate the effectiveness and efficiency of the proposed method. The predicted AZ configuration can achieve an accuracy of 89.7% within 98% of confidence level. By cause of the learning approach, the method performs even better in real scenarios, saving up to 27% of resources compared to previous work analytically modeled. I NTRODUCTION New offloading techniques to cope with the explosive growth in mobile traffic volumes, are a fundamental component of the next generation radio access network (5G). Part of the cellular traffic volume can be offloaded to vehicular-to- vehicular communication (V2V), leaving to the infrastructure the task of managing and coordinating the communication. In this context, of special interest are communication paradigms such as Floating Content (FC), an opportunistic communication scheme for the local dissemination of information [1]. FC as an infrastructure-less communication model, enables probabilistic contents storing in geographically constrained locations - denoted as Anchor Zone (AZ) - and over a limited amount of time based on the application requirements.
Label Propagation for Learning with Label Proportions
Poyiadzi, Rafael, Santos-Rodriguez, Raul, Twomey, Niall
ABSTRACT Learning with Label Proportions (LLP) is the problem of recovering the underlying true labels given a dataset when the data is presented in the form of bags. This paradigm is particularly suitable in contexts where providing individual labels is expensive and label aggregates are more easily obtained. In the healthcare domain, it is a burden for a patient to keep a detailed diary of their daily routines, but often they will be amenable to provide higher level summaries of daily behavior. We present a novel and efficient graph-based algorithm that encourages local smoothness and exploits the global structure of the data, while preserving the'mass' of each bag. 1. INTRODUCTION The whole spectrum of learning paradigms ranging from supervised to unsupervised learning is densely packed with different settings that have limited (or no) access to the true labels. For instance, in semi-supervised learning we have access to labels only for a subset of the dataset (usually small). Additionally, we can characterize the uncertainty in the labeling process when learning from noisy and partial labels [1, 2], using noise as a proxy in the former and a subset of labels per example in the later. Differently, in this paper we focus on learning from aggregated labels, also commonly referred to as Learning with Label Proportions (LLP) [3, 4]. In this setting, we assume the data comes in the form of bags of examples and for each of which we are given proportion of labels corresponding to each class.
Effective extractive summarization using frequency-filtered entity relationship graphs
Sakhadeo, Archit, Srivastava, Nisheeth
Word frequency-based methods for extractive summarization are easy to implement and yield reasonable results across languages. However, they have significant limitations - they ignore the role of context, they offer uneven coverage of topics in a document, and sometimes are disjointed and hard to read. We use a simple premise from linguistic typology - that English sentences are complete descriptors of potential interactions between entities, usually in the order subject-verb-object - to address a subset of these difficulties. We have developed a hybrid model of extractive summarization that combines word-frequency based keyword identification with information from automatically generated entity relationship graphs to select sentences for summaries. Comparative evaluation with word-frequency and topic word-based methods shows that the proposed method is competitive by conventional ROUGE standards, and yields moderately more informative summaries on average, as assessed by a large panel (N 94) of human raters.
Coarse-to-fine volumetric segmentation of teeth in Cone-Beam CT
Ezhov, Matvey, Zakirov, Adel, Gusarev, Maxim
ABSTRACT We consider the problem of localizing and segmenting individual teeth inside 3D Cone-Beam Computed Tomography (CBCT) images. To handle large image sizes we approach this task with a coarse-to-fine framework, where the whole volume is first analyzed as a 33-class semantic segmentation (adults have up to 32 teeth) in coarse resolution, followed by binary semantic segmentation of the cropped region of interest in original resolution. To improve the performance of the challenging 33-class segmentation, we first train the Coarse step model on a large weakly labeled dataset, then fine-tune it on a smaller precisely labeled dataset. The Fine step model is trained with precise labels only. Empirically, this framework yields precise teeth masks with low localization errors sufficient for many real-world applications.
Making Sense of Vision and Touch: Self-Supervised Learning of Multimodal Representations for Contact-Rich Tasks
Lee, Michelle A., Zhu, Yuke, Srinivasan, Krishnan, Shah, Parth, Savarese, Silvio, Fei-Fei, Li, Garg, Animesh, Bohg, Jeannette
Abstract-- Contact-rich manipulation tasks in unstructured environments often require both haptic and visual feedback. However, it is nontrivial to manually design a robot controller that combines modalities with very different characteristics. While deep reinforcement learning has shown success in learning control policies for high-dimensional inputs, these algorithms are generally intractable to deploy on real robots due to sample complexity. We use self-supervision to learn a compact and multimodal representation of our sensory inputs, which can then be used to improve the sample efficiency of our policy learning. We evaluate our method on a peg insertion task, generalizing over different geometry, configurations, and clearances, while being robust to external perturbations. Results for simulated and real robot experiments are presented. Even in routine tasks such as putting a car key in the ignition, humans effortlessly combine our senses of vision and touch to complete the task. Visual feedback provides information about semantic and geometric object properties for accurate reaching or grasp pre-shaping. Haptic feedback provides information about the current contact conditions between object and environment for accurate localization and control even under occlusions.
A predictive processing model of perception and action for self-other distinction
In everyday social interaction we constantly try to deduce and predict the underlying intentions behind others' social actions, like facial expressions, speech, gestures, or body posture. This is no easy problem and the underlying cognitive mechanisms and neural processes even have been dubbed the,,dark matter" of social neuroscience (Przyrembel et al., 2012). Generally, action recognition is assumed to rest upon principles of prediction-based processing (Clark, 2013), where predictions about expected sensory stimuli are continuously formed and evaluated against incoming sensory input to inform further processing. Such a predictive processing does not only inform our perception of actions of others, but also our action production in which we constantly predict the sensory consequences of our own actions and correct them in case of deviations. Both of these processes are assumed to be supported by the structure of the human sensorimotor system that is characterised by perception-action coupling (Prinz, 1997) and common coding of the underlying representations.
Spark AI Summit Europe - Developing from cloud to the edge
Organizations around the world are gearing up for a future powered by data, cloud, and Artificial Intelligence (AI). This week at Spark AI Summit Europe, I talked about how Microsoft is committed to delivering cutting-edge innovations that help our customers navigate these technological and business shifts. The driving force behind powerful AI applications is data โ and getting the most out of AI requires a modern data estate. Organizations are using their data to extract important insights to drive their businesses forward and engage their customers in ways that were simply not possible before. One such example is the Real Madrid Football Club, one of the world's top sports franchises with 500 million fans worldwide.
Artificial intelligence: Parking a car with only 12 neurons
IMAGE: This is the neural net with different layers of interconnected neurons. A naturally grown brain works quite differently than an ordinary computer program. It does not use code consisting of clear logical instructions, it is a network of cells that communicate with each other. Simulating such networks on a computer can help to solve problems which are difficult to break down into logical operations. At TU Wien (Vienna), in collaboration with researchers at Massachusetts Institute of Technology (MIT), a new approach for programming such neural networks has now been developed, which models the time evolution of the nerve signals in a completely different way.
Lyft buys an AR company to bolster its self-driving car efforts
Lyft is ramping up it self-driving car strategy on two fronts. To start, the ridesharing mainstay has acquired Blue Vision Labs, a UK-based augmented reality firm whose underlying technology helps cars both know their location and understand their surroundings. The startup will join Lyft's Level 5 team (that is, working on complete autonomy) to contribute its knowledge. TechCrunch has also learned that Blue Vision will serve as the "anchor" for a London research and development wing. This is Lyft's first buyout in the self-driving world, although it's not clear if this will be the last. "We are always evaluating build versus buy," the company's autonomous driving lead Luc Vincent told TechCrunch.