Goto

Collaborating Authors

 Statistical Learning


New Insights into Learning with Correntropy Based Regression

arXiv.org Machine Learning

Stemming from information-theoretic learning, the correntropy criterion and its applications to machine learning tasks have been extensively explored and studied. Its application to regression problems leads to the robustness enhanced regression paradigm -- namely, correntropy based regression. Having drawn a great variety of successful real-world applications, its theoretical properties have also been investigated recently in a series of studies from a statistical learning viewpoint. The resulting big picture is that correntropy based regression regresses towards the conditional mode function or the conditional mean function robustly under certain conditions. Continuing this trend and going further, in the present study, we report some new insights into this problem. First, we show that under the additive noise regression model, such a regression paradigm can be deduced from minimum distance estimation, implying that the resulting estimator is essentially a minimum distance estimator and thus possesses robustness properties. Second, we show that the regression paradigm, in fact, provides a unified approach to regression problems in that it approaches the conditional mean, the conditional mode, as well as the conditional median functions under certain conditions. Third, we present some new results when it is utilized to learn the conditional mean function by developing its error bounds and exponential convergence rates under conditional $(1+\epsilon)$-moment assumptions. The saturation effect on the established convergence rates, which was observed under $(1+\epsilon)$-moment assumptions, still occurs, indicating the inherent bias of the regression estimator. These novel insights deepen our understanding of correntropy based regression, help cement the theoretic correntropy framework, and also enable us to investigate learning schemes induced by general bounded nonconvex loss functions.


Text-Based Ideal Points

arXiv.org Machine Learning

Ideal point models analyze lawmakers' votes to quantify their political positions, or ideal points. But votes are not the only way to express a political position. Lawmakers also give speeches, release press statements, and post tweets. In this paper, we introduce the text-based ideal point model (TBIP), an unsupervised probabilistic topic model that analyzes texts to quantify the political positions of its authors. We demonstrate the TBIP with two types of politicized text data: U.S. Senate speeches and senator tweets. Though the model does not analyze their votes or political affiliations, the TBIP separates lawmakers by party, learns interpretable politicized topics, and infers ideal points close to the classical vote-based ideal points. One benefit of analyzing texts, as opposed to votes, is that the TBIP can estimate ideal points of anyone who authors political texts, including non-voting actors. To this end, we use it to study tweets from the 2020 Democratic presidential candidates. Using only the texts of their tweets, it identifies them along an interpretable progressive-to-moderate spectrum.


Demystifying Principal Component Analysis

#artificialintelligence

Data visualization has always been an essential part of any machine learning operation. It helps to get a very clear intuition about the distribution of data, which in turn helps us to decide which model is best for the problem, we are dealing with. Currently, with the advancement of machine learning, we more often need to deal with large datasets. The datasets are having a large number of features, and can only be visualized using a large feature space. Now, we can only visualize 2-dimensional planes but, visualization of data is also seems pretty necessary, as we saw in our discussion above. This is where Principal Component Analysis comes in.


20 Questions to excel in Machine Learning Interview

#artificialintelligence

Bias is error due to erroneous or overly simplistic assumptions in the learning algorithm you're using. This can lead to the model underfitting your data, making it hard for it to have high predictive accuracy and for you to generalize your knowledge from the training set to the test set. Variance is error due to too much complexity in the learning algorithm you're using. This leads to the algorithm being highly sensitive to high degrees of variation in your training data, which can lead your model to overfit the data. You'll be carrying too much noise from your training data for your model to be very useful for your test data.


Bayesian Inference: The Maximum Entropy Principle

#artificialintelligence

In this article, I will explain what the maximum entropy principle is, how to apply it and why it's useful in the context of Bayesian inference. The code to reproduce the results and figures can be found in this notebook. The maximum entropy principle is a method to create probability distributions that is most consistent with a given set of assumptions and nothing more. The rest of the article will explain what this means. First, we need to a way to measure the uncertainty in a probability distribution.


Can graph machine learning identify hate speech in online social networks?

#artificialintelligence

Over three decades, the Internet has grown from a small network of computers used by research scientists to communicate and exchange data to a technology that has penetrated almost every aspect of our day-to-day lives. Today, it is hard to imagine a life without online access for doing business, shopping, and socialising. A technology that has connected humanity at a scale never before possible has also amplified some of our worst qualities. Online hate speech spreads virally across the globe with short and long term consequences for individuals and societies. These consequences are often difficult to measure and predict. Online social media websites and mobile apps have inadvertently become the platform for the spread and proliferation of hate speech. "Hate speech is a type of speech that takes place online (e.g., the Internet, online social media platforms) with the purpose to attack a person or a group on the basis of attributes such as race, religion, ethnic origin, sexual orientation, disability, or gender."


Top 10 Machine Learning Algorithms for ML Beginners

#artificialintelligence

In the last decade machine learning becomes one of the hottest topics in the world, Andrew Ng considers it as the new electricity. In today's world machine learning powers many of the services we use -- recommendation systems like those on Netflix, YouTube, and Spotify; search engines like Google and Baidu; social-media feeds like Facebook and Twitter; voice assistants like Siri and Alexa. Having known that, let's see how machine learning works. In simple terms, machine learning algorithms use statistics to find patterns in massive amounts of data. The data are also known as the dataset, it's could contain images, texts, words, and clicks.


Deep Learning Prerequisites: Logistic Regression in Python

#artificialintelligence

Udemy Coupon - Deep Learning Prerequisites: Logistic Regression in Python, Data science techniques for professionals and students - learn the theory behind logistic regression and code in Python BESTSELLER 4.6 (2,529 ratings) Created by Lazy Programmer Inc.  English [Auto-generated], Portuguese [Auto-generated], 1 more Preview this Course - GET COUPON CODE 100% Off Udemy Coupon . Free Udemy Courses . Online Classes


Technology

#artificialintelligence

Check out their product page … link Get the Chemometrics and Spectroscopy News in real time on Twitter @ CalibModel and follow us. Near-Infrared Spectroscopy (NIRS) "Non-invasive method to identify the type of green tea inside teabag using NIR spectroscopy, support vector machines and Bayesian optimization" LINK "Online milk composition analysis with an on-farm near-infrared sensor" LINK "Anonymous fecal sampling and NIRS studies of diet quality: Problem or opportunity?" LINK "Organic and Symbiotic Fertilization of Tomato Plants Monitored by Litterbag-NIRS and Foliar-NIRS Rapid Spectroscopic Methods Running title: Litterbag-NIRS and Foliar-NIRS model in symbiotic tomato" LINK "Determination of crude protein and metabolized energy with near infrared reflectance spectroscopy (NIRS) in ruminant mixed feeds" LINK Infrared Spectroscopy (IR) and Near-Infrared Spectroscopy (NIR) "Near Infrared Spectroscopy as an efficient tool for the Qualitative and Quantitative Determination of Sugar ...


Machine Learning Tutorial

#artificialintelligence

As businesses interact with customers and collect large volumes of data, they have started appreciating the importance of machine learning in their business. By collecting insights from the data, companies can work better and gain a competitive edge over others. The Machine Learning tutorial will help you understand machine learning, it's working principles, and how it can be used every day. As an emerging field, Machine Learning offers immense opportunities for those looking at a highly impactful and satisfying career in IT. The Machine Learning market is expected to reach USD 8.81 Billion by 2022, with a growth rate of 44.1-per cent.