Goto

Collaborating Authors

 Statistical Learning


Dynamic Ensemble Selection VS K-NN: why and when Dynamic Selection obtains higher classification performance?

arXiv.org Artificial Intelligence

Multiple classifier systems focus on the combination of classifiers to obtain better performance than a single robust one. These systems unfold three major phases: pool generation, selection and integration. One of the most promising MCS approaches is Dynamic Selection (DS), which relies on finding the most competent classifier or ensemble of classifiers to predict each test sample. The majority of the DS techniques are based on the K-Nearest Neighbors (K-NN) definition, and the quality of the neighborhood has a huge impact on the performance of DS methods. In this paper, we perform an analysis comparing the classification results of DS techniques and the K-NN classifier under different conditions. Experiments are performed on 18 state-of-the-art DS techniques over 30 classification datasets and results show that DS methods present a significant boost in classification accuracy even though they use the same neighborhood as the K-NN. The reasons behind the outperformance of DS techniques over the K-NN classifier reside in the fact that DS techniques can deal with samples with a high degree of instance hardness (samples that are located close to the decision border) as opposed to the K-NN. In this paper, not only we explain why DS techniques achieve higher classification performance than the K-NN but also when DS should be used.


A Self-paced Regularization Framework for Partial-Label Learning

arXiv.org Artificial Intelligence

Partial label learning (PLL) aims to solve the problem where each training instance is associated with a set of candidate labels, one of which is the correct label. Most PLL algorithms try to disambiguate the candidate label set, by either simply treating each candidate label equally or iteratively identifying the true label. Nonetheless, existing algorithms usually treat all labels and instances equally, and the complexities of both labels and instances are not taken into consideration during the learning stage. Inspired by the successful application of self-paced learning strategy in machine learning field, we integrate the self-paced regime into the partial label learning framework and propose a novel Self-Paced Partial-Label Learning (SP-PLL) algorithm, which could control the learning process to alleviate the problem by ranking the priorities of the training examples together with their candidate labels during each learning iteration. Extensive experiments and comparisons with other baseline methods demonstrate the effectiveness and robustness of the proposed method.


QDEE: Question Difficulty and Expertise Estimation in Community Question Answering Sites

arXiv.org Artificial Intelligence

In this paper, we present a framework for Question Difficulty and Expertise Estimation (QDEE) in Community Question Answering sites (CQAs) such as Yahoo! Answers and Stack Overflow, which tackles a fundamental challenge in crowdsourcing: how to appropriately route and assign questions to users with the suitable expertise. This problem domain has been the subject of much research and includes both language-agnostic as well as language conscious solutions. We bring to bear a key language-agnostic insight: that users gain expertise and therefore tend to ask as well as answer more difficult questions over time. We use this insight within the popular competition (directed) graph model to estimate question difficulty and user expertise by identifying key hierarchical structure within said model. An important and novel contribution here is the application of "social agony" to this problem domain. Difficulty levels of newly posted questions (the cold-start problem) are estimated by using our QDEE framework and additional textual features. We also propose a model to route newly posted questions to appropriate users based on the difficulty level of the question and the expertise of the user. Extensive experiments on real world CQAs such as Yahoo! Answers and Stack Overflow data demonstrate the improved efficacy of our approach over contemporary state-of-the-art models. The QDEE framework also allows us to characterize user expertise in novel ways by identifying interesting patterns and roles played by different users in such CQAs.


From R scripts to shiny applications โ€“ use case in the spare parts business

@machinelearnbot

The main benefits of using R compared to other languages are speed and handling of large data sets. That's the cause why we started the implementation of pricing models for one of our customers in R in 2010. Today, this customer runs a shiny server with eight shiny applications for more than 500 users. Image 1 depicts the architecture of the shiny applications and internal and external interfaces. The initial R script used different data sets to run mathematical models (cluster analysis and regression models) to evaluate a fair market value for surplus spare parts.


Data Science โ€“ The New Monetization Model for Analytics Industry

@machinelearnbot

"Data Scientist is the sexiest job of the 21st century" โ€“ Harvard Business Review "Expect a shortage of over 100,000 data scientists by 2020" โ€“ Gartner Unarguably, in today's hyper-competitive marketplace, Data Science plays an indispensable role for organizations to personalize experiences and create value out of their data. Analyzing large data sets without preset defined rules or scope for analysis to uncover insights, a sublime concept till a few years ago, will form the key basis of competition in the future to significantly unlock business value, unleashing new waves of productivity for businesses, enabling a culture of innovation, and reinvigorating internal processes, as long as the right ecosystem and enablers are put in place. Numerous articles today are buzzing with this glamourous new word in the Analytics world i.e. So what exactly is Data Science or this hype around Data Scientist? Frankly speaking, multiple definitions, roles, job descriptions exist making it harder for businesses to understand what truly is the role about and the ROI out of making any additional investments.


The 10 Statistical Techniques Data Scientists Need to Master

@machinelearnbot

Regardless of where you stand on the matter of Data Science sexiness, it's simply impossible to ignore the continuing importance of data, and our ability to analyze, organize, and contextualize it. Drawing on their vast stores of employment data and employee feedback, Glassdoor ranked Data Scientist #1 in their 25 Best Jobs in America list. So the role is here to stay, but unquestionably, the specifics of what a Data Scientist does will evolve. With technologies like Machine Learning becoming ever-more common place, and emerging fields like Deep Learning gaining significant traction amongst researchers and engineers -- and the companies that hire them -- Data Scientists continue to ride the crest of an incredible wave of innovation and technological progress. While having a strong coding ability is important, data science isn't all about software engineering (in fact, have a good familiarity with Python and you're good to go).


Executing gradient descent on the earth

#artificialintelligence

A common analogy for explaining gradient descent goes like the following: a person is stuck in the mountains during heavy fog, and must navigate their way down. The natural way they will approach this is to look at the slope of the visible ground around them and slowly work their way down the mountain by following the downward slope. This captures the essence of gradient descent, but this analogy always ends up breaking down when we scale to a high dimensional space where we have very little idea what the actual geometry of that space is. Although, in the end it's often not a practical concern because gradient descent seems to work pretty well. But the important question is: how well does gradient descent perform on the actual earth? In a general model gradient descent is used to find weights for a model that minimizes our cost function, which is usually some representation of the errors made by a model over a number of predictions.


16 Free Machine Learning Books

#artificialintelligence

The following is a list of free books on Machine Learning. A Brief Introduction To Neural Networks provides a comprehensive overview of the subject of neural networks and is divided into 4 parts โ€“Part I: From Biology to Formalization -- Motivation, Philosophy, History and Realization of Neural Models,Part II: Supervised learning Network Paradigms, Part III: Unsupervised learning Network Paradigms and Part IV: Excursi, Appendices and Registers. A Course In Machine Learning is designed to provide a gentle and pedagogically organized introduction to the field and provide a view of machine learning that focuses on ideas and models, not on math. The audience of this book is anyone who knows differential calculus and discrete math, and can program reasonably well. An undergraduate in their fourth or fifth semester should be fully capable of understanding this material. However, it should also be suitable for first year graduate students, perhaps at a slightly faster pace.


10 machine learning algorithms Every Data Scientist should know in 2018

#artificialintelligence

A data scientist is a person hired to analyze and interpret complicated digital records, together with the utilization statistics of a website; particularly so that it will help an enterprise in its decision-making. An analytical model is a mathematical model that is designed to carry out a particular task or to find out the probability of a selected event i.e. the solution to the equations used to describe modifications in a system can be expressed as a mathematical analytic function. According to Layman, an analytical model is simply a mathematical presentation of an enterprise problem. A simple equation y a bx may be termed as a model with a group of predefined input data and desired output. Scalable and efficient analytical modeling is severely consequential to enable the business to use those techniques to ever-more sizably voluminous data sets for reducing the time taken to carry out these analyses. Accordingly, models are engendered that put into effect key algorithms to determine the solution to our quandary business.


Machine Learning Demystified - DZone AI

#artificialintelligence

Too often, when I hear someone talk about artificial intelligence (AI) or, more recently, machine learning (ML), the Terminator/Matrix scenario is repeated: a warning that we shouldn't meddle with powers we don't understand and that the consequences of not adhering to this warning can be dire indeed. I don't mean to say that AI/ML won't ever pose a risk in the future (though I still think we have a long way to go before the Matrix or Terminator situation), but what bothers me is that I think this fear has the telltale signs of a fear of the unknown. AI and machine learning are seen as something mysterious and therefore threatening, concepts shrouded in mystique and understood only by a few Gnostics of the AI/ML cult. I don't think it has to, or even should, be that way. So, what is machine learning?