Education
How KDnuggets Is Serving a New Generation of Data Professionals
KDnuggets has been around in one form or another for decades. It's witnessed every trend in data mining at first, and now in data science. Matthew Mayo, a Machine Learning Researcher, has served as Editor of KDnuggets for the past four years. Here, he talks to Justin Charness, Director of Product Marketing for Oracle AI, about quenching the growing thirst for knowledge about data science. The overarching trends are toward beginner-to-intermediate technical articles and tutorials on a variety of technical topics, from data science software to deep learning concepts to algorithm overviews to project implementations.
On EducationThe Data Science Course 2019: Complete Data Science - CouponED
BESTSELLER 4.5 (26,962 ratings) 122,893 students enrolled Created by 365 Careers, 365 Careers Team What you'll learn The course provides the entire toolbox you need to become a data scientist Fill up your resume with in demand data science skills: Statistical analysis, Python programming with NumPy, pandas, matplotlib, and Seaborn, Advanced statistical analysis, Tableau, Machine Learning with stats models and scikit-learn, Deep learning with TensorFlow Impress interviewers by showing an understanding of the data science field Learn how to pre-process data Understand the mathematics behind Machine Learning (an absolute must which other courses don't teach!) Start coding in Python and learn how to use it for statistical analysis Perform linear and logistic regressions in Python Carry out cluster and factor analysis Be able to create Machine Learning algorithms in Python, using NumPy, statsmodels and scikit-learn Apply your skills to real-life business cases Use state-of-the-art Deep Learning frameworks such as Google's TensorFlowDevelop a business intuition while coding and solving tasks with big data Unfold the power of deep neural networks Improve Machine Learning algorithms by studying underfitting, overfitting, training, validation, n-fold cross validation, testing, and how hyperparameters could improve performance Warm up your fingers as you will be eager to apply everything you have learned here to more and more real-life situations Requirements No prior experience is required. We will start from the very basics You'll need to install Anaconda. We will show you how to do that step by step Microsoft Excel 2003, 2010, 2013, 2016, or 365 Each of these topics builds on the previous ones. And you risk getting lost along the way if you don't acquire these skills in the right order. For example, one would struggle in the application of Machine Learning techniques before understanding the underlying Mathematics.
Will Artificial Intelligence Lead Us to a Utopian Future? Elon Musk & Jack Ma Discuss Its Prospects
The two luminaries of the technology, Alibaba's Jack Ma and Tesla's Elon Musk, shared the dias to discuss about the prospects of artificial intelligence (AI) -- its impact on humanity, education, jobs, environment, civilisation and its extinction, etc. One thing was clear from the top minds that AI will disrupt the future solving things for humans, and both seem agreed that this is not some technology that will lead to a dystopian future. While there is an ongoing debate from around the world whether AI is actually beneficial for modern society, views presented by Ma and Musk on the technology, however, were unending -- yet logical and smart. "I think AI is going to open a new chapter for the society; it will help humans understand ourselves better... I don't think AI is a threat, or terrible. People worry a lot about this today are those people that I call themโฆ uhhh'college smartiness'," Ma said.
Innovative Deep Learning Solution Developed by AI Company Sightcorp
The AI-powered software company, Sightcorp has managed to creatively iterate and improve the detection aspect of facial analysis and recognition software, thanks to their unique focus on Deep Learning, rather than the classical Haar Cascade detector methodology. This comes on the heels of an effort on the part of Sightcorp to enhance the effectiveness of their AI-powered software for users looking to gain an even deeper insight into moment-to-moment interaction. From capturing and quantifying emotions and moods to analyzing information on demographics and providing actionable and reliable data on customers' attention spans, Sightcorp intends to give users as much insight as necessary to make an informed, predictive decision. The initiative focused on deepening the software's ability to detect faces across varying head poses, with greater accuracy, speed, and granularity. The fact is that not all faces and behaviors are alike.
Best of arXiv.org for AI, Machine Learning, and Deep Learning โ July 2019 - insideBIGDATA
Researchers from all over the world contribute to this repository as a prelude to the peer review process for publication in traditional journals. We hope to save you some time by picking out articles that represent the most promise for the typical data scientist. The articles listed below represent a fraction of all articles appearing on the preprint server. They are listed in no particular order with a link to each paper along with a brief overview. Especially relevant articles are marked with a "thumbs up" icon. Consider that these are academic research papers, typically geared toward graduate students, post docs, and seasoned professionals.
An AI privacy conundrum? The neural net knows more than it says ZDNet
Artificial intelligence is the process of using a machine such as a neural network to say things about data. Most times, what is said is a simple affair, like classifying pictures into cats and dogs. Increasingly, though, AI scientists are posing questions about what the neural network "knows," if you will, that is not captured in simple goals such as classifying pictures or generating fake text and images. It turns out there's a lot left unsaid, even if computers don't really know anything in the sense a person does. Neural networks, it seems, can retain a memory of specific training data, which could open individuals whose data is captured in the training activity to violations of privacy. For example, Nicholas Carlini, formerly a student at UC Berkeley's AI lab, approached the problem of what computers "memorize" about training data, in work done with colleagues at Berkeley.
A Multimodal Alerting System for Online Class Quality Assurance
Chen, Jiahao, Li, Hang, Wang, Wenxin, Ding, Wenbiao, Huang, Gale Yan, Liu, Zitao
Online 1 on 1 class is created for more personalized learning experience. It demands a large number of teaching resources, which are scarce in China. To alleviate this problem, we build a platform (marketplace), i.e., \emph{Dahai} to allow college students from top Chinese universities to register as part-time instructors for the online 1 on 1 classes. To warn the unqualified instructors and ensure the overall education quality, we build a monitoring and alerting system by utilizing multimodal information from the online environment. Our system mainly consists of two key components: banned word detector and class quality predictor. The system performance is demonstrated both offline and online. By conducting experimental evaluation of real-world online courses, we are able to achieve 74.3\% alerting accuracy in our production environment.
Flexible Auto-weighted Local-coordinate Concept Factorization: A Robust Framework for Unsupervised Clustering
Zhang, Zhao, Zhang, Yan, Li, Sheng, Liu, Guangcan, Zeng, Dan, Yan, Shuicheng, Wang, Meng
Concept Factorization (CF) and its variants may produce inaccurate representation and clustering results due to the sensitivity to noise, hard constraint on the reconstruction error and pre-obtained approximate similarities. To improve the representation ability, a novel unsupervised Robust Flexible Auto-weighted Local-coordinate Concept Factorization (RFA-LCF) framework is proposed for clustering high-dimensional data. Specifically, RFA-LCF integrates the robust flexible CF by clean data space recovery, robust sparse local-coordinate coding and adaptive weighting into a unified model. RFA-LCF improves the representations by enhancing the robustness of CF to noise and errors, providing a flexible constraint on the reconstruction error and optimizing the locality jointly. For robust learning, RFA-LCF clearly learns a sparse projection to recover the underlying clean data space, and then the flexible CF is performed in the projected feature space. RFA-LCF also uses a L2,1-norm based flexible residue to encode the mismatch between the recovered data and its reconstruction, and uses the robust sparse local-coordinate coding to represent data using a few nearby basis concepts. For auto-weighting, RFA-LCF jointly preserves the manifold structures in the basis concept space and new coordinate space in an adaptive manner by minimizing the reconstruction errors on clean data, anchor points and coordinates. By updating the local-coordinate preserving data, basis concepts and new coordinates alternately, the representation abilities can be potentially improved. Extensive results on public databases show that RFA-LCF delivers the state-of-the-art clustering results compared with other related methods.
Leveraging Just a Few Keywords for Fine-Grained Aspect Detection Through Weakly Supervised Co-Training
Karamanolakis, Giannis, Hsu, Daniel, Gravano, Luis
User-generated reviews can be decomposed into fine-grained segments (e.g., sentences, clauses), each evaluating a different aspect of the principal entity (e.g., price, quality, appearance). Automatically detecting these aspects can be useful for both users and downstream opinion mining applications. Current supervised approaches for learning aspect classifiers require many fine-grained aspect labels, which are labor-intensive to obtain. And, unfortunately, unsupervised topic models often fail to capture the aspects of interest. In this work, we consider weakly supervised approaches for training aspect classifiers that only require the user to provide a small set of seed words (i.e., weakly positive indicators) for the aspects of interest. First, we show that current weakly supervised approaches do not effectively leverage the predictive power of seed words for aspect detection. Next, we propose a student-teacher approach that effectively leverages seed words in a bag-of-words classifier (teacher); in turn, we use the teacher to train a second model (student) that is potentially more powerful (e.g., a neural network that uses pre-trained word embeddings). Finally, we show that iterative co-training can be used to cope with noisy seed words, leading to both improved teacher and student models. Our proposed approach consistently outperforms previous weakly supervised approaches (by 14.1 absolute F1 points on average) in six different domains of product reviews and six multilingual datasets of restaurant reviews.