Europe
Digital Transformation is More Than Just a Technology Shift
Ever since its introduction in 2015, there's hardly anyone who denies information regularly on the big'Digital Transformation' that India boasts about. The efforts are being made to turn the Digital India a reality in the coming few years. On the other hand, recent advancements in key technologies have enhanced IoT's capacities and utility within enterprises and supported growth through digital transformation. BIS Infotech's, Jyoti Gazmer, recently had a dialogue with Interra Information Technologies Aniruddha Guha Sarkar, Senior Vice President of Engineering, Interra Information Technologies & Ranjan Guha, President North America, Interra Information Technologies has brought analysis of the digital transformation in India-what we know so far, how do we measure it and the challenges. Also, it touches majorly on how IoT can be an extended definition to define Digital Transformations.
Machine Learning Best Algorithms: Gradient Boosting Machines (GBM)
We'll have a main talk (30 mins) and 3 excellent lightning talks about the machine learning algorithm that usually achieves the best accuracy on structured/tabular data (e.g. in industry/business applications or in Kaggle competitions): Abstract: With all the hype about deep learning and "AI", it is not well publicized that for structured/tabular data widely encountered in business applications it is actually another machine learning algorithm, the gradient boosting machine (GBM) that most often achieves the highest accuracy in supervised learning tasks. In this talk we'll review some of the main GBM implementations available as R and Python packages such as xgboost, h2o, lightgbm etc, we'll discuss some of their main features and characteristics, and we'll see how tuning GBMs and creating ensembles of the best models can achieve the best prediction accuracy for many business problems. Bio: Szilard studied Physics in the 90s and obtained a PhD by using statistical methods to analyze the risk of financial portfolios. He worked in finance, then more than a decade ago moved to become the Chief Scientist of a tech company in Santa Monica doing everything data (analysis, modeling, data visualization, machine learning, data infrastructure etc). He is the founder/organizer of several meetups in the Los Angeles area (R, data science etc) and the data science community website datascience.la.
Inside SAP's Data Analytics Startup
When Helen Arnold agreed to move 6,000 miles from Germany to Mountain View to head up SAP Data Network, the longtime SAP executive faced some big challenges. Not only was she bootstrapping a data analytics software firm to help Fortune 1000 clients digitally transform themselves – a sizable hurdle in its own right -- but she needed to address big cultural differences between Silicon Valley and Walldorf. SAP Data Network opened for business in May 2017 with the goal of leveraging a mix of SAP software like SAP Data Hub, Leonardo, and HANA, along with open source data science tools and professional services, to help clients build their own transformational data products. By standardizing on reusable cloud-based software components – and even reusable data feeds, where available -- the company seeks to help clients build innovative data products in a matter of weeks or months, rather than the years it often takes for organizations that are starting from scratch. So far, the startup has worked with about 40 clients, most of which are companies that use SAP's well-regarded ERP suite.
Bayesian Optimization of Combinatorial Structures
Baptista, Ricardo, Poloczek, Matthias
The optimization of expensive-to-evaluate black-box functions over combinatorial structures is an ubiquitous task in machine learning, engineering and the natural sciences. The combinatorial explosion of the search space and costly evaluations pose challenges for current techniques in discrete optimization and machine learning, and critically require new algorithmic ideas (NIPS BayesOpt 2017). This article proposes, to the best of our knowledge, the first algorithm to overcome these challenges, based on an adaptive, scalable model that identifies useful combinatorial structure even when data is scarce. Our acquisition function pioneers the use of semidefinite programming to achieve efficiency and scalability. Experimental evaluations demonstrate that this algorithm consistently outperforms other methods from combinatorial and Bayesian optimization.
Domain Adaptation for Infection Prediction from Symptoms Based on Data from Different Study Designs and Contexts
Rehman, Nabeel Abdur, Aliapoulios, Maxwell Matthaios, Umarwani, Disha, Chunara, Rumi
Acute respiratory infections have epidemic and pandemic potential and thus are being studied worldwide, albeit in many different contexts and study formats. Predicting infection from symptom data is critical, though using symptom data from varied studies in aggregate is challenging because the data is collected in different ways. Accordingly, different symptom profiles could be more predictive in certain studies, or even symptoms of the same name could have different meanings in different contexts. We assess state-of-the-art transfer learning methods for improving prediction of infection from symptom data in multiple types of health care data ranging from clinical, to home-visit as well as crowdsourced studies. We show interesting characteristics regarding six different study types and their feature domains. Further, we demonstrate that it is possible to use data collected from one study to predict infection in another, at close to or better than using a single dataset for prediction on itself. We also investigate in which conditions specific transfer learning and domain adaptation methods may perform better on symptom data. This work has the potential for broad applicability as we show how it is possible to transfer learning from one public health study design to another, and data collected from one study may be used for prediction of labels for another, even collected through different study designs, populations and contexts.
Variational Bi-domain Triplet Autoencoder
Kuznetsova, Rita, Bakhteev, Oleg
We investigate deep generative models, which allow us to use training data from one domain to build a model for another domain. We consider domains to have similar structure (texts, images). We propose the Variational Bi-domain Triplet Autoencoder (VBTA) that learns a joint distribution of objects from different domains. There are many cases when obtaining any supervision (e.g. paired data) is difficult or ambiguous. For such cases we can seek a method that is able to the information about data relation and structure from the latent space. We extend the VBTAs objective function by the relative constraints or triplets that sampled from the shared latent space across domains. In other words, we combine the deep generative model with a metric learning ideas in order to improve the final objective with the triplets information. We demonstrate the performance of the VBTA model on different tasks: bi-directional image generation, image-to-image translation, even on unpaired data. We also provide the qualitative analysis. We show that VBTA model is comparable and outperforms some of the existing generative models.
Smart Inverter Grid Probing for Learning Loads: Part II - Probing Injection Design
Bhela, Siddharth, Kekatos, Vassilis, Veeramachaneni, Sriharsha
This two-part work puts forth the idea of engaging power electronics to probe an electric grid to infer non-metered loads. Probing can be accomplished by commanding inverters to perturb their power injections and record the induced voltage response. Once a probing setup is deemed topologically observable by the tests of Part I, Part II provides a methodology for designing probing injections abiding by inverter and network constraints to improve load estimates. The task is challenging since system estimates depend on both probing injections and unknown loads in an implicit nonlinear fashion. The methodology first constructs a library of candidate probing vectors by sampling over the feasible set of inverter injections. Leveraging a linearized grid model and a robust approach, the candidate probing vectors violating voltage constraints for any anticipated load value are subsequently rejected. Among the qualified candidates, the design finally identifies the probing vectors yielding the most diverse system states. The probing task under noisy phasor and non-phasor data is tackled using a semidefinite-program (SDP) relaxation. Numerical tests using synthetic and real-world data on a benchmark feeder validate the conditions of Part I; the SDP-based solver; the importance of probing design; and the effects of probing duration and noise.
Diffusion Scattering Transforms on Graphs
Gama, Fernando, Ribeiro, Alejandro, Bruna, Joan
Stability is a key aspect of data analysis. In many applications, the natural notion of stability is geometric, as illustrated for example in computer vision. Scattering transforms construct deep convolutional representations which are certified stable to input deformations. This stability to deformations can be interpreted as stability with respect to changes in the metric structure of the domain. In this work, we show that scattering transforms can be generalized to non-Euclidean domains using diffusion wavelets, while preserving a notion of stability with respect to metric changes in the domain, measured with diffusion maps. The resulting representation is stable to metric perturbations of the domain while being able to capture "high-frequency" information, akin to the Euclidean Scattering.
Forecasting Internally Displaced Population Migration Patterns in Syria and Yemen
Huynh, Benjamin Q., Basu, Sanjay
Armed conflict has led to an unprecedented number of internally displaced persons (IDPs) - individuals who are forced out of their homes but remain within their country. IDPs often urgently require shelter, food, and healthcare, yet prediction of when large fluxes of IDPs will cross into an area remains a major challenge for aid delivery organizations. Accurate forecasting of IDP migration would empower humanitarian aid groups to more effectively allocate resources during conflicts. We show that monthly flow of IDPs from province to province in both Syria and Yemen can be accurately forecasted one month in advance, using publicly available data. We model monthly IDP flow using data on food price, fuel price, wage, geospatial, and news data. We find that machine learning approaches can more accurately forecast migration trends than baseline persistence models. Our findings thus potentially enable proactive aid allocation for IDPs in anticipation of forecasted arrivals.
Learning Qualitatively Diverse and Interpretable Rules for Classification
Ross, Andrew Slavin, Pan, Weiwei, Doshi-Velez, Finale
There has been growing interest in developing accurate models that can also be explained to humans. Unfortunately, if there exist multiple distinct but accurate models for some dataset, current machine learning methods are unlikely to find them: standard techniques will likely recover a complex model that combines them. In this work, we introduce a way to identify a maximal set of distinct but accurate models for a dataset. We demonstrate empirically that, in situations where the data supports multiple accurate classifiers, we tend to recover simpler, more interpretable classifiers rather than more complex ones.