Europe
50 Shades of Grey – The Psychology of a Data Scientist
Unless you've recently graduated from one of the new Data Science courses that have been popping up online and in various universities around the world, then becoming a Data Scientist was most likely slightly accidental and was more about the journey than the destination. I started out as a physicist and had a strong mathematical grounding, but I had a passion for medicine. After completing my bachelor's degree I took a master's degree in medical physics. This is where I gained an appreciation for the importance of image analysis and the role that data plays in medicine. I created a virtual model of a human torso by segmenting images from the Visible Human Project.
Changing landscape of a digitised world: Are you ready?
We are living in a society where change is exponential - perhaps it has always been. What types of technologies will be foundational to our business? The disruption of Blockchain to data and databases parallels the disruption of the Internet to communications and networks. From Wall Street to Fintech Accelerators, incredible investments at Blockchain have led to incredible speed of innovation. This disruption is now ready to penetrate enterprise IT.
Internet of Things: Are We There Yet? (The 2016 IoT Landscape)
Is the Internet of Things the world's most confusing tech trend? On the one hand, we're told it's going to be epic, and soon – all predictions are either in tens of billions (of connected devices) and trillions (of dollars of economic value to be created). On the other hand, the dominant feeling expressed by end users (including at this year's CES show, arguably the bellwether of the industry) is essentially "meh" – right now the IoT feels like an avalanche of new connected products, many of which seem to solve trivial, "first world" problems: expensive gadgets that resolutely fall in the "nice to have" category, rather than "must have". And, for all the talk about a mega tech trend, things seem to be moving at the speed of molasses, with little discernible progress year on year. Part of the problem is perhaps one of semantics. While gadgets are indeed part of the category (and quite often very large markets onto themselves), the Internet of Things (which we define as any "connected hardware" other than desktops, laptops and smartphones) is a much broader, and deeper, trend that cuts across both the consumer, enterprise and industrial spaces.
Whole-brain substitute CT generation using Markov random field mixture models
Hildeman, Anders, Bolin, David, Wallin, Jonas, Johansson, Adam, Nyholm, Tufve, Asklund, Thomas, Yu, Jun
Computed tomography (CT) equivalent information is needed for attenuation correction in PET imaging and for dose planning in radiotherapy. Prior work has shown that Gaussian mixture models can be used to generate a substitute CT (s-CT) image from a specific set of MRI modalities. This work introduces a more flexible class of mixture models for s-CT generation, that incorporates spatial dependency in the data through a Markov random field prior on the latent field of class memberships associated with a mixture model. Furthermore, the mixture distributions are extended from Gaussian to normal inverse Gaussian (NIG), allowing heavier tails and skewness. The amount of data needed to train a model for s-CT generation is of the order of 100 million voxels. The computational efficiency of the parameter estimation and prediction methods are hence paramount, especially when spatial dependency is included in the models. A stochastic Expectation Maximization (EM) gradient algorithm is proposed in order to tackle this challenge. The advantages of the spatial model and NIG distributions are evaluated with a cross-validation study based on data from 14 patients. The study show that the proposed model enhances the predictive quality of the s-CT images by reducing the mean absolute error with 17.9%. Also, the distribution of CT values conditioned on the MR images are better explained by the proposed model as evaluated using continuous ranked probability scores.
Approachability of convex sets in generalized quitting games
Flesch, János, Laraki, Rida, Perchet, Vianney
We consider Blackwell approachability, a very powerful and geometric tool in game theory, used for example to design strategies of the uninformed player in repeated games with incomplete information. We extend this theory to "generalized quitting games" , a class of repeated stochastic games in which each player may have quitting actions, such as the Big-Match. We provide three simple geometric and strongly related conditions for the weak approachability of a convex target set. The first is sufficient: it guarantees that, for any fixed horizon, a player has a strategy ensuring that the expected time-average payoff vector converges to the target set as horizon goes to infinity. The third is necessary: if it is not satisfied, the opponent can weakly exclude the target set. In the special case where only the approaching player can quit the game (Big-Match of type I), the three conditions are equivalent and coincide with Blackwell's condition. Consequently, we obtain a full characterization and prove that the game is weakly determined-every convex set is either weakly approachable or weakly excludable. In games where only the opponent can quit (Big-Match of type II), none of our conditions is both sufficient and necessary for weak approachability. We provide a continuous time sufficient condition using techniques coming from differential games, and show its usefulness in practice, in the spirit of Vieille's seminal work for weak approachability.Finally, we study uniform approachability where the strategy should not depend on the horizon and demonstrate that, in contrast with classical Blackwell approacha-bility for convex sets, weak approachability does not imply uniform approachability.
Predictive Coarse-Graining
Schöberl, Markus, Zabaras, Nicholas, Koutsourelakis, Phaedon-Stelios
We propose a data-driven, coarse-graining formulation in the context of equilibrium statistical mechanics. In contrast to existing techniques which are based on a fine-to-coarse map, we adopt the opposite strategy by prescribing a probabilistic coarse-to-fine map. This corresponds to a directed probabilistic model where the coarse variables play the role of latent generators of the fine scale (all-atom) data. From an information-theoretic perspective, the framework proposed provides an improvement upon the relative entropy method and is capable of quantifying the uncertainty due to the information loss that unavoidably takes place during the CG process. Furthermore, it can be readily extended to a fully Bayesian model where various sources of uncertainties are reflected in the posterior of the model parameters. The latter can be used to produce not only point estimates of fine-scale reconstructions or macroscopic observables, but more importantly, predictive posterior distributions on these quantities. Predictive posterior distributions reflect the confidence of the model as a function of the amount of data and the level of coarse-graining. The issues of model complexity and model selection are seamlessly addressed by employing a hierarchical prior that favors the discovery of sparse solutions, revealing the most prominent features in the coarse-grained model. A flexible and parallelizable Monte Carlo - Expectation-Maximization (MC-EM) scheme is proposed for carrying out inference and learning tasks. A comparative assessment of the proposed methodology is presented for a lattice spin system and the SPC/E water model.
StruClus: Structural Clustering of Large-Scale Graph Databases
We present a structural clustering algorithm for large-scale datasets of small labeled graphs, utilizing a frequent subgraph sampling strategy. A set of representatives provides an intuitive description of each cluster, supports the clustering process, and helps to interpret the clustering results. The projection-based nature of the clustering approach allows us to bypass dimensionality and feature extraction problems that arise in the context of graph datasets reduced to pairwise distances or feature vectors. While achieving high quality and (human) interpretable clusterings, the runtime of the algorithm only grows linearly with the number of graphs. Furthermore, the approach is easy to parallelize and therefore suitable for very large datasets. Our extensive experimental evaluation on synthetic and real world datasets demonstrates the superiority of our approach over existing structural and subspace clustering algorithms, both, from a runtime and quality point of view.
Statistical and computational trade-offs in estimation of sparse principal components
Wang, Tengyao, Berthet, Quentin, Samworth, Richard J.
In recent years, sparse principal component analysis has emerged as an extremely popular dimension reduction technique for high-dimensional data. The theoretical challenge, in the simplest case, is to estimate the leading eigenvector of a population covariance matrix under the assumption that this eigenvector is sparse. An impressive range of estimators have been proposed; some of these are fast to compute, while others are known to achieve the minimax optimal rate over certain Gaussian or sub-Gaussian classes. In this paper, we show that, under a widely-believed assumption from computational complexity theory, there is a fundamental trade-off between statistical and computational performance in this problem. More precisely, working with new, larger classes satisfying a restricted covariance concentration condition, we show that there is an effective sample size regime in which no randomised polynomial time algorithm can achieve the minimax optimal rate. We also study the theoretical performance of a (polynomial time) variant of the well-known semidefinite relaxation estimator, revealing a subtle interplay between statistical and computational efficiency.
Narcissists may start out popular, but people see through them in the long run
But if, as they say in this electoral season, you're looking to "grow your base," exercising emotional intelligence -- expressing empathy, checking your emotions in a bid to avoid conflict, and investing in personal relationships -- is a strategy that beats narcissism over the long term. A new exploration of how we make friends and influence people rigorously measured the emergence of popularity in small groups -- first-year college students organized into 15 study groups of about 20 in Poland. In the first week of their assignment to a group and then again three months later, 170 of the freshmen named the person or people they most liked in their group. Upon recruitment into the study, each participant completed standard inventories assessing their narcissistic personality traits and gauging their emotional intelligence. The findings: When a group of strangers is thrown together, individuals who score high on narcissism enjoy an early surge of admiration, recognition and friendship among their peers.
Toward human-centric A.I.
Twenty years ago, Stuart Russell co-wrote a book titled Artificial Intelligence: A Modern Approach (AIMA), destined to become the dominant text in its field. Near the end of the book, he posed a question: "What if A.I. does succeed?" Today, progress toward human-level artificial intelligence (A.I.) is advancing rapidly, and Russell, a professor of computer science, is posing the same question with more urgency. The benefits of A.I. are not at issue. If improperly constrained, Russell warns, a machine as smart as or smarter than humans "is of no use whatsoever -- in fact it's catastrophic."