Technology
Clustering Time-Series Energy Data from Smart Meters
Lavin, Alexander, Klabjan, Diego
Investigations have been performed into using clustering methods in data mining time-series data from smart meters. The problem is to identify patterns and trends in energy usage profiles of commercial and industrial customers over 24-hour periods, and group similar profiles. We tested our method on energy usage data provided by several U.S. power utilities. The results show accurate grouping of accounts similar in their energy usage patterns, and potential for the method to be utilized in energy efficiency programs.
New metrics for learning and inference on sets, ontologies, and functions
Yang, Ruiyu, Jiang, Yuxiang, Hahn, Matthew W., Housworth, Elizabeth A., Radivojac, Predrag
We propose new metrics on sets, ontologies, and functions that can be used in various stages of probabilistic modeling, including exploratory data analysis, learning, inference, and result interpretation. These new functions unify and generalize some of the popular metrics on sets and functions, such as the Jaccard and bag distances on sets and Marczewski-Steinhaus distance on functions. We then introduce information-theoretic metrics on directed acyclic graphs drawn independently according to a fixed probability distribution and show how they can be used to calculate similarity between class labels for the objects with hierarchical output spaces (e.g., protein function). Finally, we provide evidence that the proposed metrics are useful by clustering species based solely on functional annotations available for subsets of their genes. The functional trees resemble evolutionary trees obtained by the phylogenetic analysis of their genomes.
A Convergent Gradient Descent Algorithm for Rank Minimization and Semidefinite Programming from Random Linear Measurements
Zheng, Qinqing, Lafferty, John
Semidefinite programming has become a key optimization tool in many areas of applied mathematics, signal processing and machine learning. SDPs often arise naturally from the problem structure, or are derived as surrogate optimizations that are relaxations of difficult combinatorial problems [7, 1, 8]. In spite of the importance of SDPs in principle--promising efficient algorithms with polynomial runtime guarantees--it is widely recognized that current optimization algorithms based on interior point methods can handle only relatively small problems. Thus, a considerable gap exists between the theory and applicability of SDP formulations. Scalable algorithms for semidefinite programming, and closely related families of nonconvex programs more generally, are greatly needed. A parallel development is the surprising effectiveness of simple classical procedures such as gradient descent for large scale problems, as explored in the recent machine learning literature. In many areas of machine learning and signal processing such as classification, deep learning, and phase retrieval, gradient descent methods, in particular first order stochastic optimization, have led to remarkably efficient algorithms that can attack very large scale problems [3, 2, 10, 6]. In this paper we build on this work to develop first-order algorithms for solving the rank minimization problem under random measurements and a closely related family of semidefinite programs. Our algorithms are efficient and scalable, and we prove that they attain linear convergence to the global optimum under natural assumptions.
Solution Template for Energy Demand Forecasting
The post is by Ilan Reiter, Principal Data Science Manager at Microsoft. The past few years have witnessed dramatic changes to the energy sector. Renewable energy sources along with the emergence of IoT (Internet of Things) are creating exciting new opportunities. On the consumption side, utilities and indeed the entire energy sector have seen consumption flatten out, with consumers demanding better ways to monitor and control their energy usage. Furthermore, with many grids becoming outdated and expensive to maintain, utilities and smart grid companies are in ever greater need to innovate.
How Machine Learning Works [Interactive]
As an employee of a company that provides digital security through machine learning, Stephanie Yee spent a lot of time familiarizing clients with the secret sauce behind her product. So she and her colleague, designer Tony Chu, set out to create an interactive graphic that would do the explaining for them. The pair chose a topic they thought would be intuitive to most people--real estate prices--and created an interactive environment that builds in complexity as the user scrolls. In the first 30 days, the site got 250,000 page views worldwide. Feedback showed Chu and Yee that experts in many fields could use their interactives.
SparkR (R on Spark) - Spark 1.6.0 Documentation
SparkR is an R package that provides a light-weight frontend to use Apache Spark from R. In Spark 1.6.0, SparkR provides a distributed data frame implementation that supports operations like selection, filtering, aggregation etc. (similar to R data frames, dplyr) but on large datasets. SparkR also supports distributed machine learning using MLlib. A DataFrame is a distributed collection of data organized into named columns. It is conceptually equivalent to a table in a relational database or a data frame in R, but with richer optimizations under the hood.
China's Baidu Releases Its AI Code
Google and Facebook aren't the only ones vying to be the standard bearer for the hottest AI technique around. China's leading Internet search company, Baidu, which is also investing heavily in a popular and powerful machine-learning technology called deep learning, today released some key code that it uses to make this AI software run very efficiently. Baidu's code was recently used to build an impressive speech-recognition system called Deep Speech 2. For some short sentences, this system is better than most humans at recognizing speech correctly (see "Baidu's Deep-Learning System Rivals People at Speech Recognition"). This is an especially useful technology for Baidu, because it offers a better way for the company's many millions of users to access its services, especially on mobile. Typing Chinese characters on a smartphone is tricky and complex, and many people in China already prefer to use their voice to send short messages or to search the Web for information.
Automation may mean a post-work society but we shouldn't be afraid
When researchers Frey and Osborne predicted in 2013 that 47% of US jobs were susceptible to automation by 2050, they set off a wave of dystopian concern. But the key word is "susceptible". The automation revolution is possible, but without a radical change in the social conventions surrounding work it will not happen. The real dystopia is that, fearing the mass unemployment and psychological aimlessness it might bring, we stall the third industrial revolution. Instead we end up creating millions of low skilled jobs that do not need to exist.