mlib
Introduction to Machine Learning Interviews Book · MLIB
You can read the web-friendly version of the book here. You can find the source code on GitHub. The Discord to discuss the answers to the questions in the book is here. As a candidate, I've interviewed at a dozen big companies and startups. I've got offers for machine learning roles at companies including Google, NVIDIA, Snap, Netflix, Primer AI, and Snorkel AI. I've also been rejected at many other companies.
Chapter 7. Machine learning workflows · MLIB
Even though deep learning seems to be all that people in the research community is talking about, most real-world problems are still being solved by classical machine learning algorithms including k-nearest neighbor and XGBoost. In this chapter, we will cover fundamentals that are essential for understanding machine learning algorithms, as well as non-deep learning algorithms that you might find useful in both your day-to-day jobs and interviews.
State of Machine Learning with Apache Spark
Apache Spark is now a household name for Machine Learning and distributed Data Science. This has made Spark the go-to choice for big data machine learning applications even if it requires work with terabytes or petabytes of data. It's efficient and very convenient usage model and interface has led to many adopting this as their prime platform for distributed machine learning. Also, the Apache Spark works as a dominating force for data analysts. The twenty-first century has been the story of resources getting smaller and cheaper while accessibility to these devices has reached higher numbers each quarter.
How Apache Spark Became Essential For Machine Learning
Spark library has a library for ML labelled as MLib. This Apache Spark library has algorithms for the functions of classification, regression, clustering, collaborative filtering, dimensionality reduction, etc. The classification includes classifying things into different categories. For example, in emails the classification is done in categories of inbox, sent, drafts, spam and so on. Clustering example is bifurcating the news on the basis of the title and content of the news.
Google's Machine Learning Buy, Alteryx Updates: Big Data Roundup - InformationWeek
Google acquires another machine learning startup, predictive analytics and wearables are driving the continued adoption of electronic health records (EHR), millennials are believers in the value of data-driven projects, Alteryx has released a new set of capabilities for its analytics platform, and H2O.ai has released an updated API for Apache Spark. The self-service data analytics company has added a new set of enhancements to the Alteryx Analytics platform. Among them is the ability to create models that progress from descriptive to predictive to prescriptive analytics. Alteryx said that this lets users evaluate possible alternatives and predict outcomes through simulation analysis. Other new capabilities include predictive modeling support for users working inside Teradata data warehouses with in-database analytics.
Google's Machine Learning Buy, Alteryx Updates: Big Data Roundup - InformationWeek
Google acquires another machine learning startup, predictive analytics and wearables are driving the continued adoption of electronic health records (EHR), Millennials are believers in the value of data-driven projects, Alteryx has released a new set of capabilities for its analytics platform, and H2O.ai has released an updated API for Apache Spark. The self-service data analytics company has added a new set of enhancements to the Alteryx Analytics platform. Among them is the ability to create models that progress from descriptive to predictive to prescriptive analytics. Alteryx said that this lets users evaluate possible alternatives and predict outcomes through simulation analysis. Other new capabilities include predictive modeling support for users working inside Teradata data warehouses with in-database analytics.