Country
Do Compressed Representations Generalize Better?
Hafez-Kolahi, Hassan, Kasaei, Shohreh, Soleymani-Baghshah, Mahdiyeh
One of the most studied problems in machine learning is finding reasonable constraints that guarantee the generalization of a learning algorithm. These constraints are usually expressed as some simplicity assumptions on the target. For instance, in the V apnik-Chervonenkis (VC) theory the space of possible hypotheses is considered to have a limited VC dimension and in kernel methods there are assumptions on the spectrum of the operator in the Hilbert space. One way to formulate the simplicity assumption is via information theoretic concepts. In this paper, the constraint on the entropy H ( X) of the input variable X is studied as a simplicity assumption. It is proven that the sample complexity to achieve an null -δ Probably Approximately Correct (P AC) hypothesis is bounded by 2 2 H ( X) / null log 1 δ null 2 which is sharp up to the 1 null 2 factor. Morever, it is shown that if a feature learning process is employed to learn the compressed representation from the dataset, this bound no longer exists. These findings have important implications on the Information Bottleneck (IB) theory which had been utilized to explain the generalization power of Deep Neural Networks (DNNs), but its applicability for this purpose is currently under debate by researchers. In particular, this is a rigorous proof for the previous heuristic that compressed representations are expnentially easier to be learned. However, our analysis pinpoints two factors preventing the IB, in its current form, to be applicable in studying neural networks. Firstly, the exponential dependence of sample complexity on 1 / null, which can lead to a dramatic e ff ect on the bounds in practical applications when null is small. Secondly, our analysis reveals that arguments based on input compression are inherently insu fficient to explain generalization of methods like DNNs in which the features are also learned using available data. Keywords: Compressed Representation; Generalization Bound; Information Bottleneck. 1. Introduction The main objective of learning is to develop algorithms which can learn general patterns by using a finite number of samples drawn from a target distribution. The "no free lunch" theorem states that if there is no constraint on the distribution, it is impossible to say anything about the samples not seen in the training set (Wolpert, 1996b).
A Layered Architecture for Active Perception: Image Classification using Deep Reinforcement Learning
Mousavi, Hossein K., Liu, Guangyi, Yuan, Weihang, Takáč, Martin, Muñoz-Avila, Héctor, Motee, Nader
We propose a planning and perception mechanism for a robot (agent), that can only observe the underlying environment partially, in order to solve an image classification problem. A three-layer architecture is suggested that consists of a meta-layer that decides the intermediate goals, an action-layer that selects local actions as the agent navigates towards a goal, and a classification-layer that evaluates the reward and makes a prediction. We design and implement these layers using deep reinforcement learning. A generalized policy gradient algorithm is utilized to learn the parameters of these layers to maximize the expected reward. Our proposed methodology is tested on the MNIST dataset of handwritten digits, which provides us with a level of explainability while interpreting the agent's intermediate goals and course of action.
Gradual Network for Single Image De-raining
Huang, Zhe, Yu, Weijiang, Zhang, Wayne, Feng, Litong, Xiao, Nong
Most advances in single image de-raining meet a key challenge, which is removing rain streaks with different scales and shapes while preserving image details. Existing single image de-raining approaches treat rain-streak removal as a process of pixel-wise regression directly. However, they are lacking in mining the balance between over-de-raining (e.g. removing texture details in rain-free regions) and under-de-raining (e.g. leaving rain streaks). In this paper, we firstly propose a coarse-to-fine network called Gradual Network (GraNet) consisting of coarse stage and fine stage for delving into single image de-raining with different granularities. Specifically, to reveal coarse-grained rain-streak characteristics (e.g. long and thick rain streaks/raindrops), we propose a coarse stage by utilizing local-global spatial dependencies via a local-global subnetwork composed of region-aware blocks. Taking the residual result (the coarse de-rained result) between the rainy image sample (i.e. the input data) and the output of coarse stage (i.e. the learnt rain mask) as input, the fine stage continues to de-rain by removing the fine-grained rain streaks (e.g. light rain streaks and water mist) to get a rain-free and well-reconstructed output image via a unified contextual merging sub-network with dense blocks and a merging block. Solid and comprehensive experiments on synthetic and real data demonstrate that our GraNet can significantly outperform the state-of-the-art methods by removing rain streaks with various densities, scales and shapes while keeping the image details of rain-free regions well-preserved.
Understanding and Robustifying Differentiable Architecture Search
Zela, Arber, Elsken, Thomas, Saikia, Tonmoy, Marrakchi, Yassine, Brox, Thomas, Hutter, Frank
Differentiable Architecture Search (DARTS) has attracted a lot of attention due to its simplicity and small search costs achieved by a continuous relaxation and an approximation of the resulting bi-level optimization problem. However, DARTS does not work robustly for new problems: we identify a wide range of search spaces for which DARTS yields degenerate architectures with very poor test performance. We study this failure mode and show that, while DARTS successfully minimizes validation loss, the found solutions generalize poorly when they coincide with high validation loss curvature in the space of architectures. We show that by adding one of various types of regularization we can robustify DARTS to find solutions with smaller Hessian spectrum and with better generalization properties. Based on these observations we propose several simple variations of DARTS that perform substantially more robustly in practice. Our observations are robust across five search spaces on three image classification tasks and also hold for the very different domains of disparity estimation (a dense regression task) and language modelling. We provide our implementation and scripts to facilitate reproducibility.
Repositioning Bikes with Carrier Vehicles and Bike Trailers in Bike Sharing Systems
Zheng, Xinghua, Tang, Ming, Zhuo, Hankz Hankui, Wen, Kevin X.
Bike Sharing Systems (BSSs) have been adopted in many major cities of the world due to traffic congestion and carbon emissions. Although there have been approaches to exploiting either bike trailers via crowdsourcing or carrier vehicles to reposition bikes in the "right" stations in the "right" time, they do not jointly consider the usage of both bike trailers and carrier vehicles. In this paper, we aim to take advantage of both bike trailers and carrier vehicles to reduce the loss of demand with regard to the crowdsourcing of bike trailers and the fuel cost of carrier vehicles. In the experiment, we exhibit that our approach outperforms baselines in several datasets from bike sharing companies. Introduction Bike sharing systems (BSSs) typically have a set of base stations that are strategically placed throughout a city and each station has a fixed number of docks, e.g., Capital Bike-share 1, Bluebikes 2, Mobike 3, BIXI 4, etc. At the beginning of the day, each station is stocked with a predetermined number of bikes. Customers can pick and drop bikes from any station and are charged depending on the hiring duration (Tsai, Chen, and Hong 2019; Hulot, Aloise, and Jena 2018; Lowalekar et al. 2017; Vulcano, van Ryzin, and Ratliff 2012; Schuijbroek, Hampshire, and van Hoeve 2017). Due to the individualistic and uncoordinated movements of customers, there is often starvation (empty base stations precluding bike pickup) or congestion (full base stations precluding bike return) of bikes at certain stations, which results in a significant loss of customer demand (Shu et al. 2013; Chen, Liu, and Liu 2018). To address this problem, a variety of systems (Ghosh et al. 2017; Lowalekar et al. 2017) employ the idea of repositioning idle bikes with the help of carrier vehicles during the day, by taking into account the movement of bikes by customers (Tsai, Chen, and Hong 2019; Pfrommer et al. 2014; Ghosh and V arakantham 2017).
AIBA: An AI Model for Behavior Arbitration in Autonomous Driving
Trasnea, Bogdan, pozna, Claudiu, Grigorescu, Sorin
Driving in dynamically changing traffic is a highly challenging task for autonomous vehicles, especially in crowded urban roadways. The Artificial Intelligence (AI) system of a driverless car must be able to arbitrate between different driving strategies in order to properly plan the car's path, based on an understandable traffic scene model. In this paper, an AI behavior arbitration algorithm for Autonomous Driving (AD) is proposed. The method, coined AIBA (AI Behavior Arbitration), has been developed in two stages: (i) human driving scene description and understanding and (ii) formal modelling. The description of the scene is achieved by mimicking a human cognition model, while the modelling part is based on a formal representation which approximates the human driver understanding process. The advantage of the formal representation is that the functional safety of the system can be analytically inferred. The performance of the algorithm has been evaluated in Virtual Test Drive (VTD), a comprehensive traffic simulator, and in GridSim, a vehicle kinematics engine for prototypes.
Hybrid Probabilistic Inference with Logical Constraints: Tractability and Message-Passing
Zeng, Zhe, Yan, Fanqi, Morettin, Paolo, Vergari, Antonio, Broeck, Guy Van den
Weighted model integration (WMI) is a very appealing framework for probabilistic inference: it allows to express the complex dependencies of real-world hybrid scenarios where variables are heterogeneous in nature (both continuous and discrete) via the language of Satisfiability Modulo Theories (SMT); as well as computing probabilistic queries with arbitrarily complex logical constraints. Recent work has shown WMI inference to be reducible to a model integration (MI) problem, under some assumptions, thus effectively allowing hybrid probabilistic reasoning by volume computations. In this paper, we introduce a novel formulation of MI via a message passing scheme that allows to efficiently compute the marginal densities and statistical moments of all the variables in linear time. As such, we are able to amortize inference for arbitrarily rich MI queries when they conform to the problem structure, here represented as the primal graph associated to the SMT formula. Furthermore, we theoretically trace the tractability boundaries of exact MI. Indeed, we prove that in terms of the structural requirements on the primal graph that make our MI algorithm tractable - bounding its diameter and treewidth - the bounds are not only sufficient, but necessary for tractable inference via MI.
New Omnitracs Exec Takes the Wheel of its Data, AI, and Machine Learning Operations » Dallas Innovates
All parts of the transportation industry are being affected by next-gen technology with the AI market in transportation estimated to grow by nearly 20 percent annually to $10.3 billion by 2030. To support a strategy to enhance its products with more data-driven and artificial intelligence/machine learning-based solutions, Dallas-based Omnitracs LLC--a leading provider of fleet management solutions to transportation and logistics companies.--has In this role, Bose--who has a Ph.D. in artificial intelligence--will oversee Omnitracs' data and AI offerings and operations, helping its customers with actionable insights that support their broader business goals, according to a statement. "Omnitracs has a deep history of providing progressive technologies to the transportation and logistics industry," Bose said in a statement. "Digital transformation is reshaping the industry rapidly, and I look forward to leveraging Omnitracs' rich data securely and applying data science and AI techniques to bring new value to Omnitracs customers."
Machine Learning for IT
Driverless AI is H2O.ai's latest flagship product for automatic machine learning. It fully automates some of the most challenging and productive tasks in applied data science such as feature engineering, model tuning, model ensembling and production deployment. Driverless AI turns Kaggle-winning grandmaster recipes into production-ready code (Java and C), and is specifically designed to avoid common mistakes such as under- or overfitting, data leakage or improper model validation, which are some of the hardest challenges in data science. Other industry-leading capabilities include automatic data visualization and machine learning interpretability. We're now excited to add the ability for users, partners and customers to extend the platform with Bring-Your-Own-Recipe.
Fury as AI app gives racist labels and calls people a 'rape suspect'
A viral app which claims to'honestly' classify selfies using its in-built artificial intelligence has been spewing out vile and racist labels. Furious users say their pictures have been slapped with offensive and racist terms such as'negro', 'slant eye' and'rape suspect' by the app which was developed at Princeton University. Developers say causing offence was exactly the intention and it was intended to be deliberately provocative to draw attention to the in-built prejudice and discrimination in many forms of machine learning. But many users are still furious that their images have played a seemingly unwitting part in the controversial project. One MailOnline staffer who tried the app was dubbed a'rape suspect' when he uploaded his selfie.