Goto

Collaborating Authors

 Genre


What's Driving Apache Spark Growth? SQL, Streaming and Machine Learning -- ADTmag

#artificialintelligence

Databricks Inc., the primary commercial steward behind the popular open source Apache Spark data processing framework for Big Data analytics, published a new report indicating the technology is still red-hot, driven by more use of SQL, streaming analytics and machine learning. The company this summer polled more than 900 organizations and solicited data from 1,615 respondents -- mostly Spark users -- coming from the ranks of data scientists, data engineers, architects and others, and last week published the results in the Apache Spark Survey 2016 Report (free download upon providing registration info). The report follows up on a similar survey last year, confirming the technology's widespread popularity as the most active open source project in the Big Data space. "As in 2015, which was a tremendous year in growth for Apache Spark, this year, too, its growth remains unabated -- not only in areas like the public cloud, but also with the increased use of Spark Streaming and the use of machine learning," the report states. "2016 also shows Spark's robust adoption across a variety of organizations and users from many functional roles to build complex solutions, using multiple Spark components."


Tim Cook Tells Investors Apple Is Investing Heavily in Machine Learning R&D

#artificialintelligence

Apple revealed its fourth quarter financial results during a conference call with investors on Tuesday, but perhaps the most interesting narrative -- aside from the company's first annual revenue decline since 2001 -- was the tech giants avowed focus on machine learning. "Today, machine learning drives improvement in countless features across our products," Apple CEO Tim Cook boasted during his initial remarks. Cook then offered a state of machine learning at Apple, by running down areas where machine learning is already helping to improve the iPhone user experience, noting that machine learning "enables the proactive features in iOS 10, and that cameras employ it in face recognizing software. Machine learning is also a key aspect of Apple's fitness offerings. "Machine learning continually helps Siri get smarter in areas including understanding natural language," Cook added.


5 Free Statistics eBooks You Need to Read This Autumn

#artificialintelligence

I hope you enjoy them, and it would be great if you would leave brief reviews of these books in the comments below โ€“ I'm sure all the authors would appreciate your comments and shares. About the Author Lee Baker is an award-winning software creator with a passion for turning data into a story. A proud Yorkshireman, he now lives by the sparkling shores of the East Coast of Scotland. Physicist, statistician and programmer, child of the flower-power psychedelic '60s, it's amazing he turned out so normal! Turning his back on a promising academic career to do something more satisfying, as the CEO and co-founder of Chi-Squared Innovations he now works double the hours for half the pay and 10 times the stress - but 100 times the fun! He also wanted to be rich, famous and good looking.


Microsoft makes its deep learning tools available to all

Engadget

The same internal, deep learning tools that Microsoft engineers used to build its human-like speech recognition engine, as well as consumer products like Skype Translator and Cortana, are now available for public use. Redmond announced today that it is open-sourcing the Cognitive Toolkit that has led to many key developments coming out of its dedicated AI division. In other words: anyone can now train their own artificial intelligence. Formerly known as the CNTK, Microsoft says the beta version of the Cognitive Toolkit is not only faster than previous incarnations, but it is also beats out competing deep learning toolkits โ€“ especially when crunching large datasets across multiple machines. On a more practical level for startups and hobbyists, Microsoft says the platform is flexible enough to run on a solo laptop -- just in case you don't have a server farm loaded with NVIDIA GPUs at your disposal.


Free thinking

BBC News

A university without any teachers has opened in California this month. It's called 42 - the name taken from the answer to the meaning of life, from the science fiction series The Hitchhiker's Guide to the Galaxy. The US college, a branch of an institution in France with the same name, will train about a thousand students a year in coding and software development by getting them to help each other with projects, then mark one another's work. This might seem like the blind leading the blind - and it's hard to imagine parents at an open day being impressed by a university offering zero contact hours. But since 42 started in Paris in 2013, applications have been hugely oversubscribed. Recent graduates are now working at companies including IBM, Amazon, and Tesla, as well as starting their own firms.


Game changer

BBC News

It was the global gaming craze of the summer with players taking to the streets to try to catch on-screen monsters like Pikachu and Snorlax in real-world locations. Reports about the possible health benefits of Pokemon Go have until now been largely anecdotal, but new research suggests that playing the augmented reality game significantly increases users' activity levels regardless of their age, sex or weight, and could even extend life expectancy if kept up indefinitely. The study by researchers at Stanford University and Microsoft Research in the US analysed movement data shared by 32,000 users of the wearable device Microsoft Band, and web queries on search engine Bing over a three month period from the date the game launched in the US. Their findings show that the average Pokemon Go player took 192 more steps per day for each of the 30 days after they started playing, rising to 1,473 extra steps being taken by highly engaged players - about 25 percent more than before they started playing the game. The researchers estimate that Pokemon Go has added a total of 144 billion steps to physical activity in the US over the period of their study and that the game has been able to increase physical activity in men and women of all ages, weights, and prior activity levels.


Clarifai raises $30M to give developers visual search capabilities

#artificialintelligence

Matt Zeiler grew up in a Canadian farming community -- but fast forward a few decades and he's now running a startup that's looking to bring the same kinds of visual search tools that Pinterest and Google have to other companies and developers. That company is Clarifai, a New York-based startup that offers developers the ability to tag metadata to photos in such a way that the company algorithmically learns what kinds of objects are in photos. With that, Clarifai developers can train algorithms to be able to search for those objects, or input their own photos in order to find similar objects. The company said today that it has raised $30 million. The round was led by Menlo Ventures, with Union Square Ventures, Lux Capital and others participating.


Estimating the Size of a Large Network and its Communities from a Random Sample

arXiv.org Machine Learning

Most real-world networks are too large to be measured or studied directly and there is substantial interest in estimating global network properties from smaller sub-samples. One of the most important global properties is the number of vertices/nodes in the network. Estimating the number of vertices in a large network is a major challenge in computer science, epidemiology, demography, and intelligence analysis. In this paper we consider a population random graph G = (V;E) from the stochastic block model (SBM) with K communities/blocks. A sample is obtained by randomly choosing a subset W and letting G(W) be the induced subgraph in G of the vertices in W. In addition to G(W), we observe the total degree of each sampled vertex and its block membership. Given this partial information, we propose an efficient PopULation Size Estimation algorithm, called PULSE, that correctly estimates the size of the whole population as well as the size of each community. To support our theoretical analysis, we perform an exhaustive set of experiments to study the effects of sample size, K, and SBM model parameters on the accuracy of the estimates. The experimental results also demonstrate that PULSE significantly outperforms a widely-used method called the network scale-up estimator in a wide variety of scenarios. We conclude with extensions and directions for future work.


Recurrent switching linear dynamical systems

arXiv.org Machine Learning

Many natural systems, such as neurons firing in the brain or basketball teams traversing a court, give rise to time series data with complex, nonlinear dynamics. We can gain insight into these systems by decomposing the data into segments that are each explained by simpler dynamic units. Building on switching linear dynamical systems (SLDS), we present a new model class that not only discovers these dynamical units, but also explains how their switching behavior depends on observations or continuous latent states. These "recurrent" switching linear dynamical systems provide further insight by discovering the conditions under which each unit is deployed, something that traditional SLDS models fail to do. We leverage recent algorithmic advances in approximate inference to make Bayesian inference in these models easy, fast, and scalable.


Bayesian latent structure discovery from multi-neuron recordings

arXiv.org Machine Learning

Neural circuits contain heterogeneous groups of neurons that differ in type, location, connectivity, and basic response properties. However, traditional methods for dimensionality reduction and clustering are ill-suited to recovering the structure underlying the organization of neural circuits. In particular, they do not take advantage of the rich temporal dependencies in multi-neuron recordings and fail to account for the noise in neural spike trains. Here we describe new tools for inferring latent structure from simultaneously recorded spike train data using a hierarchical extension of a multi-neuron point process model commonly known as the generalized linear model (GLM). Our approach combines the GLM with flexible graph-theoretic priors governing the relationship between latent features and neural connectivity patterns. Fully Bayesian inference via P\'olya-gamma augmentation of the resulting model allows us to classify neurons and infer latent dimensions of circuit organization from correlated spike trains. We demonstrate the effectiveness of our method with applications to synthetic data and multi-neuron recordings in primate retina, revealing latent patterns of neural types and locations from spike trains alone.