Materials
The learning curve: From the Internet to Big Data to IoT - Industrial Internet Now
Mikko Marsio, Vice President of Digital Business and IoT at Empower group, says that what has unfolded over the past two decades and led companies to where they are today can be understood as both an evolution from a technological perspective, as well as a revolution from an industry and business perspective. From the speculative nature of the IT bubble, to the profoundness of the Internet of Things, Marsio explains how consolidating technology with business is now more imperative than ever before. "I remember a prediction that was made before I attended an MIT Executive Education course on the Internet in 2000. It envisioned the Internet becoming like electricity, meaning something that we don't even acknowledge when using," Marsio reminisces. "If you look at what was laid out in 2000 in conjunction with the IT bubble – for example that the best years for the pulp and paper industry were then and there – no one could actually have predicted how many paper mills would be shut down over the following 15 years. In order for these mills to stay relevant, they must adapt what they are producing. Companies in general need to understand how both digitalization and end-users are causing their businesses to change. Over the past few years, increasingly many have come to recognize this," he continues.
Mining Process Model Descriptions of Daily Life through Event Abstraction
Tax, Niek, Sidorova, Natalia, Haakma, Reinder, van der Aalst, Wil M. P.
Process mining techniques focus on extracting insight in processes from event logs. Process mining has the potential to provide valuable insights in (un)healthy habits and to contribute to ambient assisted living solutions when applied on data from smart home environments. However, events recorded in smart home environments are on the level of sensor triggers, at which process discovery algorithms produce overgeneralizing process models that allow for too much behavior and that are difficult to interpret for human experts. We show that abstracting the events to a higher-level interpretation can enable discovery of more precise and more comprehensible models. We present a framework for the extraction of features that can be used for abstraction with supervised learning methods that is based on the XES IEEE standard for event logs. This framework can automatically abstract sensor-level events to their interpretation at the human activity level, after training it on training data for which both the sensor and human activity events are known. We demonstrate our abstraction framework on three real-life smart home event logs and show that the process models that can be discovered after abstraction are more precise indeed.
Text Mining 101: Mining Information From A Resume
This article demonstrates a framework for mining relevant entities from a text resume. It shows how separation of parsing logic from entity specification can be achieved. Although only one resume sample is considered here, the framework can be enhanced further to be used not only for different resume formats, but also for documents such as judgments, contracts, patents, medical papers, etc. Majority of world's unstructured data is in the textual form. To make sense of it, one must, either go through it painstakingly or employ certain automated techniques to extract relevant information. Looking at the volume, variety and velocity of such textual data, it is imperative to employ Text Mining techniques to extract the relevant information, transforming unstructured data into structured form, so that further insights, processing, analysis, visualizations are possible.
Applying Machine Learning to Text Mining with Amazon S3 and RapidMiner
By some estimates, 80% of an organization's data is unstructured content. This content includes web pages, call center transcripts, surveys, feedback forms, legal documents, forums, social media, and blog articles. Therefore, organizations must analyze not just transactional information but also textual content to gain insight and boost performance. A powerful way to analyze this textual content is by using text mining. Text mining typically applies machine learning techniques such as clustering, classification, association rules and predictive modeling.
Data Mining vs. Statistics vs. Machine Learning
Data science is solely based on data. If your data is good you will get good results else, you might have heard of famous data science proverb – Garbage in Garbage out. A good (rather useful I should say) data science product is like a recipe even if one ingredient is not good, final product will not amuse the audience. If you would like more information about Data Science Training, click the Request Info. Assuming you understand your business requirement sufficiently, let's discuss what do we mean by Data Mining, Statistics & Machine Learning?
The coal miner who became a data miner
A heavy maintenance superintendent for a surface coal mine in Elgin, Texas, Evans was responsible for figuring out how to patch or replace outdated parts of a field delivery system that ferried coal from the mine to a plant. Each minute of downtime could cost the company as much as $170. Now the third-generation coal miner gets her adrenaline rush sitting indoors on a soft swivel chair, fixing code on a computer screen. The 33-year-old is a data scientist currently doing a paid residency at Galvanize in Austin. "I was an adrenaline junkie," sad Evans of her past career.
Data Mining Techniques: For Marketing, Sales, and Customer Relationship Management: Gordon S. Linoff, Michael J. A. Berry: 9780470650936: Amazon.com: Books
Who will remain a loyal customer and who won't? Which messages are most effective with which segments? How can customer value be maximized? This book supplies powerful tools for extracting the answers to these and other crucial business questions from the corporate databases where they lie buried. In the years since the first edition of this book, data mining has grown to become an indispensable tool of modern business.
Phil Libin exits General Catalyst for All Turtles, a new AI 'startup studio'
AI is one of the buzzwords of the moment in the world of tech, with startups coming at the concept from all angles -- computer vision, machine learning, unstructured data inference and natural language processing being just a handful -- in a wider effort to create more intelligent machines. Now comes a new organization that hopes to find and foster the next wave of AI businesses and products, co-founded by the ex-CEO of Evernote, Phil Libin (pictured above), who has left his role as a managing director at General Catalyst to build it (but he tells me he'll stay on as an advisor). All Turtles, as the new company is called, is not your traditional startup incubator. In an interview with TechCrunch earlier, Libin (whose other co-founders are Jessica Collier (Product Design) and Jon Cifuentes (Research and Operations) described it as "startup studio", more akin to Netflix's push to develop original content than to 500 Startups. It will start out with locations in San Francisco, Tokyo and Paris.
Machine Learning Molecular Dynamics for the Simulation of Infrared Spectra
Gastegger, Michael, Behler, Jörg, Marquetand, Philipp
Machine learning has emerged as an invaluable tool in many research areas. In the present work, we harness this power to predict highly accurate molecular infrared spectra with unprecedented computational efficiency. To account for vibrational anharmonic and dynamical effects -- typically neglected by conventional quantum chemistry approaches -- we base our machine learning strategy on ab initio molecular dynamics simulations. While these simulations are usually extremely time consuming even for small molecules, we overcome these limitations by leveraging the power of a variety of machine learning techniques, not only accelerating simulations by several orders of magnitude, but also greatly extending the size of systems that can be treated. To this end, we develop a molecular dipole moment model based on environment dependent neural network charges and combine it with the neural network potentials of Behler and Parrinello. Contrary to the prevalent big data philosophy, we are able to obtain very accurate machine learning models for the prediction of infrared spectra based on only a few hundreds of electronic structure reference points. This is made possible through the introduction of a fully automated sampling scheme and the use of molecular forces during neural network potential training. We demonstrate the power of our machine learning approach by applying it to model the infrared spectra of a methanol molecule, n-alkanes containing up to 200 atoms and the protonated alanine tripeptide, which at the same time represents the first application of machine learning techniques to simulate the dynamics of a peptide. In all these case studies we find excellent agreement between the infrared spectra predicted via machine learning models and the respective theoretical and experimental spectra.