Goto

Collaborating Authors

 Performance Analysis


Salient Object Detection: A Survey

arXiv.org Artificial Intelligence

Detecting and segmenting salient objects in natural scenes, often referred to as salient object detection, has attracted a lot of interest in computer vision. While many models have been proposed and several applications have emerged, yet a deep understanding of achievements and issues is lacking. We aim to provide a comprehensive review of the recent progress in salient object detection and situate this field among other closely related areas such as generic scene segmentation, object proposal generation, and saliency for fixation prediction. Covering 228 publications, we survey i) roots, key concepts, and tasks, ii) core techniques and main modeling trends, and iii) datasets and evaluation metrics in salient object detection. We also discuss open problems such as evaluation metrics and dataset bias in model performance and suggest future research directions.


How Do Machine Learning Programs "Learn"?

#artificialintelligence

In this article, we look at two machine learning (ML) techniques, Naive Bayes classifier and neural networks, and demystify how they work. With all the hype surrounding self-driving cars and video-game-playing AI robots, it's worth taking a step back and reminding ourselves how machine learning programs actually "learn". In this article, we look at two machine learning (ML) techniquesโ€“spam filters and neural networksโ€“and demystify how they work. And if you're not sure what machine learning even is, read about the difference between artificial intelligence, machine learning, and deep learning. One common machine learning algorithm is the Naive Bayes classifier, which is used for filtering spam emails.


Using a Customized Cost Function to Deal With Unbalanced Data - DZone AI

#artificialintelligence

As pointed out in this KDnuggets article, we often only have a few examples of the thing that we want to predict in our data. The use cases are countless: only a small part of our website visitors purchase eventually, only a few of our transactions are fraudulent, etc. This is a real problem when using machine learning. That's because the algorithms usually need many examples of each class to extract the general rules in your data, and the instances in minority classes can be discarded as noise, causing some useful rules to never be found. The KDnuggets article explained several techniques that can be used to address this problem.


As final number emergers, showtime calls Mayweather-McGregor "massive" pay-per-view success

Los Angeles Times

The "one-time-only" boxing match between a 40-year-old who retired two years ago and an Irishman making his pro debut in the sport is positioned to become the greatest-selling pay-per-view fight of all time Friday. Showtime Executive Vice President Stephen Espinoza said "it's too early to declare a hard number" but Saturday's Floyd Mayweather Jr.-Conor McGregor fight is "tracking in the mid-to-high 4 million pay-per view buys." "If we don't reach the record, we're going to be very, very close," and "we consider it a massive success." "It was an exciting, entertaining fight and there was massive interest," in it, Espinoza told The Times, crediting strong digital sales to boost the overall domestic sales. The bout is also expected to surpass the $600 million generated in total revenue by Mayweather's less-entertaining unanimous-decision triumph over seven-division champion Manny Pacquiao, with final pay-per-view numbers expected by next week.


Cross- Validation Code Visualization: Kind of Fun โ€“ Towards Data Science โ€“ Medium

@machinelearnbot

As the name of the suggests, cross-validation is the next fun thing after learning Linear Regression because it helps to improve your prediction using the K-Fold strategy. What is K-Fold you asked? Everything is explained below with Code. We are copying the target in dataset to y variable. To see the dataset uncomment the print line.


Using a Customized Cost Function to deal with Unbalanced Data

#artificialintelligence

As pointed in this Kdnuggets article, it's often the case that we only have a few examples of the thing that we want to predict in our data. The use cases are countless: only a small part of our website visitors purchase eventually, only a few of our transactions are fraudulent, etc. This is a real problem when using Machine Learning. That's because the algorithms usually need many examples of each class to extract the general rules in your data, and the instances in minority classes can be discarded as noise, causing some useful rules to never be found. The Kdnuggets article explained several techniques that can be used to address this problem.


Statistics For Data Scientist Review - Data Science Consulting

#artificialintelligence

This is great, in the sense that you don't have to worry about accidently forgetting to carry the 1 or remember how each rule in calculus operates. It is still great to have a general understanding of some of the equations you can utilize, distributions you can model and general statistics rules that can help clean up your data! We need to quickly lay out some definitions. In this post we will talk about discrete variables. If you have not heard the term before this references variables that are of a limited set. It actually could include numbers that are decimals pending on the set of variables you are using. However, these rules need to be established. For instance, you can't have 3.5783123 medical procedures in real life.


Continual One-Shot Learning of Hidden Spike-Patterns with Neural Network Simulation Expansion and STDP Convergence Predictions

arXiv.org Machine Learning

This paper presents a constructive algorithm that achieves successful one-shot learning of hidden spike-patterns in a competitive detection task. It has previously been shown (Masquelier et al., 2008) that spike-timing-dependent plasticity (STDP) and lateral inhibition can result in neurons competitively tuned to repeating spike-patterns concealed in high rates of overall presynaptic activity. One-shot construction of neurons with synapse weights calculated as estimates of converged STDP outcomes results in immediate selective detection of hidden spike-patterns. The capability of continual learning is demonstrated through the successful one-shot detection of new sets of spike-patterns introduced after long intervals in the simulation time. Simulation expansion (Lightheart et al., 2013) has been proposed as an approach to the development of constructive algorithms that are compatible with simulations of biological neural networks. A simulation of a biological neural network may have orders of magnitude fewer neurons and connections than the related biological neural systems; therefore, simulated neural networks can be assumed to be a subset of a larger neural system. The constructive algorithm is developed using simulation expansion concepts to perform an operation equivalent to the exchange of neurons between the simulation and the larger hypothetical neural system. The dynamic selection of neurons to simulate within a larger neural system (hypothetical or stored in memory) may be a starting point for a wide range of developments and applications in machine learning and the simulation of biology.


Significance testing in non-sparse high-dimensional linear models

arXiv.org Machine Learning

In high-dimensional linear models, the sparsity assumption is typically made, stating that most of the parameters are equal to zero. Under the sparsity assumption, estimation and, recently, inference have been well studied. However, in practice, sparsity assumption is not checkable and more importantly is often violated, with a large number of covariates expected to be associated with the response, indicating that possibly all, rather than just a few, parameters are non-zero. A natural example is a genome-wide gene expression profiling, where all genes are believed to affect a common disease marker. We show that existing inferential methods are sensitive to the sparsity assumption, and may, in turn, result in the severe lack of control of Type-I error. In this article, we propose a new inferential method, named CorrT, which is robust to model misspecification and adaptive to the sparsity assumption. CorrT is shown to have Type I error approaching the nominal level for \textit{any} models and Type II error approaching zero for sparse and many dense models. In fact, CorrT is also shown to be optimal in a variety of frameworks: sparse, non-sparse and hybrid models where sparse and dense signals are mixed. Numerical experiments show a favorable performance of the CorrT test compared to the state-of-the-art methods.


Stem-ming the Tide: Predicting STEM attrition using student transcript data

arXiv.org Machine Learning

Science, technology, engineering, and math (STEM) fields play growing roles in national and international economies by driving innovation and generating high salary jobs. Yet, the US is lagging behind other highly industrialized nations in terms of STEM education and training. Furthermore, many economic forecasts predict a rising shortage of domestic STEM-trained professions in the US for years to come. One potential solution to this deficit is to decrease the rates at which students leave STEM-related fields in higher education, as currently over half of all students intending to graduate with a STEM degree eventually attrite. However, little quantitative research at scale has looked at causes of STEM attrition, let alone the use of machine learning to examine how well this phenomenon can be predicted. In this paper, we detail our efforts to model and predict dropout from STEM fields using one of the largest known datasets used for research on students at a traditional campus setting. Our results suggest that attrition from STEM fields can be accurately predicted with data that is routinely collected at universities using only information on students' first academic year. We also propose a method to model student STEM intentions for each academic term to better understand the timing of STEM attrition events. We believe these results show great promise in using machine learning to improve STEM retention in traditional and non-traditional campus settings.