Genre
Born for it
Nathan Ensmenger is a professor at Indiana University who has specialised in the social and historical aspects of computing. In his book "The Computer Boys Take Over", he explores the origins of our profession, and how programmers were first hired and trained: Little has yet been written about the silent majority of computer specialists, the vast armies of largely anonymous engineers, analysts, and programmers who designed and constructed the complex systems that make possible our increasingly computerized society. The title of the book is a reference to where it all started: With the "Computer Girls". The women programming the ENIAC -- one of the very first electronic, general purpose, digital computers -- are widely considered to be the first programmers. At the time, the word "programmer", or the concept of a program, did not even exist yet.
VelocityChess Selects Zoomi to Revolutionize Online Gaming Education
To coincide with the announcement, Zoomi released details on the success of VelocityChess' education portal since implementing Zoomi's technology platform last January. The course, which includes a series of videos, presentations and infographics, caters to a wide variety of skill levels – from novice to chessmaster. Using predictive analytics and machine learning, the Zoomi platform adapts course content in real-time based on an initial diagnostic quiz, as well as data collected as the user moves through the course. Zoomi's proprietary algorithms and learning analytics allow VelocityChess to interpret the user's behavioral, performance and social patterns, adapt the content to their skill level, and even predict whether or not they'll get the next question right. To date, more than 1,400 users have taken the course.
How the Moth Radio Hour helped scientists map out meaning in the brain
This is your brain on stories. By tracking the blood flow in people's brains as they listened to a storytelling radio show, scientists at UC Berkeley have mapped out where the meanings associated with basic words are encoded in the cortex, creating the first semantic atlas of the brain. The findings, described in the journal Nature, provide an unprecedented view of language and meaning as it plays out on our neural terrain, and could potentially offer a road map for those looking to help patients with certain types of aphasia or other neurological disorders. For a long time, researchers thought about language as a primarily left-hemisphere function that took place in specific spots of the brain, such as Broca's area and Wernicke's area. But those areas aren't associated with understanding language but producing it – speech, in short.
Semantic Visualization with Neighborhood Graph Regularization
Visualization of high-dimensional data, such as text documents, is useful to map out the similarities among various data points. In the high-dimensional space, documents are commonly represented as bags of words, with dimensionality equal to the vocabulary size. Classical approaches to document visualization directly reduce this into visualizable two or three dimensions. Recent approaches consider an intermediate representation in topic space, between word space and visualization space, which preserves the semantics by topic modeling. While aiming for a good fit between the model parameters and the observed data, previous approaches have not considered the local consistency among data instances. We consider the problem of semantic visualization by jointly modeling topics and visualization on the intrinsic document manifold, modeled using a neighborhood graph. Each document has both a topic distribution and visualization coordinate. Specifically, we propose an unsupervised probabilistic model, called Semafore, which aims to preserve the manifold in the lower-dimensional spaces through a neighborhood regularization framework designed for the semantic visualization task. To validate the efficacy of Semafore, our comprehensive experiments on a number of real-life text datasets of news articles and Web pages show that the proposed methods outperform the state-of-the-art baselines on objective evaluation metrics.
Exploiting Causality for Selective Belief Filtering in Dynamic Bayesian Networks
Albrecht, Stefano V., Ramamoorthy, Subramanian
Dynamic Bayesian networks (DBNs) are a general model for stochastic processes with partially observed states. Belief filtering in DBNs is the task of inferring the belief state (i.e. the probability distribution over process states) based on incomplete and noisy observations. This can be a hard problem in complex processes with large state spaces. In this article, we explore the idea of accelerating the filtering task by automatically exploiting causality in the process. We consider a specific type of causal relation, called passivity, which pertains to how state variables cause changes in other variables. We present the Passivity-based Selective Belief Filtering (PSBF) method, which maintains a factored belief representation and exploits passivity to perform selective updates over the belief factors. PSBF produces exact belief states under certain assumptions and approximate belief states otherwise, where the approximation error is bounded by the degree of uncertainty in the process. We show empirically, in synthetic processes with varying sizes and degrees of passivity, that PSBF is faster than several alternative methods while achieving competitive accuracy. Furthermore, we demonstrate how passivity occurs naturally in a complex system such as a multi-robot warehouse, and how PSBF can exploit this to accelerate the filtering task.
A Probabilistic Adaptive Search System for Exploring the Face Space
Abad, Andres G., Castro, Luis I. Reyes
Face recall is a basic human cognitive process performed routinely, e.g., when meeting someone and determining if we have met that person before. Assisting a subject during face recall by suggesting candidate faces can be challenging. One of the reasons is that the search space - the face space - is quite large and lacks structure. A commercial application of face recall is facial composite systems - such as Identikit, PhotoFIT, and CD-FIT - where a witness searches for an image of a face that resembles his memory of a particular offender. The inherent uncertainty and cost in the evaluation of the objective function, the large size and lack of structure of the search space, and the unavailability of the gradient concept makes this problem inappropriate for traditional optimization methods. In this paper we propose a novel evolutionary approach for searching the face space that can be used as a facial composite system. The approach is inspired by methods of Bayesian optimization and differs from other applications in the use of the skew-normal distribution as its acquisition function. This choice of acquisition function provides greater granularity, with regularized, conservative, and realistic results.
Two Differentially Private Rating Collection Mechanisms for Recommender Systems
Recommender Systems (RS) [1] are a kind of system that seek to recommend to users what they are likely interested in. Unlike search engines, the users do not need to type any keyword. The RS's will learn their interest automatically. For instance, if the user has just bought a numeric camera, the RS will recommend to him some SD memory cards; if a user watches a lot of action movies, the RS may suggest some other action movies to him. And this is the typical behaviors which we observe universally in Netflix (movies), Youtube (videos), Google Play (apps), Facebook (friends), Amazon (goods) and other platforms today.
Sequential Bayesian optimal experimental design via approximate dynamic programming
Huan, Xun, Marzouk, Youssef M.
The design of multiple experiments is commonly undertaken via suboptimal strategies, such as batch (open-loop) design that omits feedback or greedy (myopic) design that does not account for future effects. This paper introduces new strategies for the optimal design of sequential experiments. First, we rigorously formulate the general sequential optimal experimental design (sOED) problem as a dynamic program. Batch and greedy designs are shown to result from special cases of this formulation. We then focus on sOED for parameter inference, adopting a Bayesian formulation with an information theoretic design objective. To make the problem tractable, we develop new numerical approaches for nonlinear design with continuous parameter, design, and observation spaces. We approximate the optimal policy by using backward induction with regression to construct and refine value function approximations in the dynamic program. The proposed algorithm iteratively generates trajectories via exploration and exploitation to improve approximation accuracy in frequently visited regions of the state space. Numerical results are verified against analytical solutions in a linear-Gaussian setting. Advantages over batch and greedy design are then demonstrated on a nonlinear source inversion problem where we seek an optimal policy for sequential sensing.
Robust subspace recovery by Tyler's M-estimator
This paper considers the problem of robust subspace recovery: given a set of $N$ points in $\mathbb{R}^D$, if many lie in a $d$-dimensional subspace, then can we recover the underlying subspace? We show that Tyler's M-estimator can be used to recover the underlying subspace, if the percentage of the inliers is larger than $d/D$ and the data points lie in general position. Empirically, Tyler's M-estimator compares favorably with other convex subspace recovery algorithms in both simulations and experiments on real data sets.
Women In Machine Learning: Katie Malone Udacity
For resources, the single best thing you can do is find people who can challenge you and make you think. These can be collaborators that you work with in "real life," or folks online (say, for example, contributing to open source projects). I've also found that the projects that turn out the best for me are the ones that I find most interesting or exciting, so I've grown to put a lot of effort into reading about many different things so I can find out what seems most cool or fun and then go after that--at first it felt a little backward, like instead I should be reading up to find out what I "should" be excited about and then letting that guide my choices, but I've found that thinking about it instead from the perspective of "what makes me excited, and let's think of a way to apply machine learning or data science to that" is way more fun for me. That's not really a resource, sorry, but I think it's important. For resources, I love online courses (like Udacity of course, but there are lots of good ones out there), podcasts (I have to say that, since I host one as a side project–Linear Digressions), and there are some excellent blogs out there too.