Goto

Collaborating Authors

 Inductive Learning


Handling imbalanced dataset in supervised learning using family of SMOTE algorithm.

#artificialintelligence

The algorithm adaptively updates the distribution and there are no assumptions made for the underlying distribution of the data. The algorithm uses Euclidean distance for KNN Algorithm. The key difference between ADASYN and SMOTE is that the former uses a density distribution, as a criterion to automatically decide the number of synthetic samples that must be generated for each minority sample by adaptively changing the weights of the different minority samples to compensate for the skewed distributions. The latter generates the same number of synthetic samples for each original minority sample.


Sergey Levine: Deep Robotic Learning CMU RI Seminar

Robohub

Abstract: "Deep learning methods have provided us with remarkably powerful, flexible, and robust solutions in a wide range of passive perception areas: computer vision, speech recognition, and natural language processing. However, active decision making domains such as robotic control present a number of additional challenges, standard supervised learning methods do not extend readily to robotic decision making, where supervision is difficult to obtain. In this talk, I will discuss experimental results that hint at the potential of deep learning to transform robotic decision making and control, present a number of algorithms and models that can allow us to combine expressive, high-capacity deep models with reinforcement learning and optimal control, and describe some of our recent work on scaling up robotic learning through collective learning with multiple robots."


A Neural Probabilistic Structured-Prediction Method for Transition-Based Natural Language Processing

Journal of Artificial Intelligence Research

We propose a neural probabilistic structured-prediction method for transition-based natural language processing, which integrates beam search and contrastive learning. The method uses a global optimization model, which can leverage arbitrary features over non-local context. Beam search is used for efficient heuristic decoding, and contrastive learning is performed for adjusting the model according to search errors. When evaluated on both chunking and dependency parsing tasks, the proposed method achieves significant accuracy improvements over the locally normalized greedy baseline on the two tasks, respectively.


Just What Is Deep Learning, and What Does It Solve In Marketing?

#artificialintelligence

How is a network trained? When given input data with a labeled answer (meaning it's already been classified), the deviation between the network's predictions and the actual answer produces an error signal. The error signal and non-linearity of each neuron's decision function tells us whether to increase or decrease each weight. The error signal gets propagated backwards all the way to the lowest layer. Over many training examples, the network weights are repeatedly tuned until finally reaching some satisfactory benchmark, such as accuracy level.


Using Graphs of Classifiers to Impose Declarative Constraints on Semi-supervised Learning

arXiv.org Machine Learning

We propose a general approach to modeling semi-supervised learning (SSL) algorithms. Specifically, we present a declarative language for modeling both traditional supervised classification tasks and many SSL heuristics, including both well-known heuristics such as co-training and novel domain-specific heuristics. In addition to representing individual SSL heuristics, we show that multiple heuristics can be automatically combined using Bayesian optimization methods. We experiment with two classes of tasks, link-based text classification and relation extraction. We show modest improvements on well-studied link-based classification benchmarks, and state-of-the-art results on relation-extraction tasks for two realistic domains.


Article 1: Why Machine Learning? โ€“ Apurba Learns ML

#artificialintelligence

For those who aren't familiar with the show, it depicts an all-seeing AI that can predict crime and other immoral acts before they even happen and passes on the information to the Government. One of the episodes of the show illustrates the origin of the Machine, where its creator -- Harold Finch -- is teaching it to distinguish between good and bad by showing it "examples". This is a perfect instance of Supervised learning -- we have a data-set or example set with the right answers given. Afterwards, we expect the computer to predict things based on the data-set. Now, to be a bit more specific, the above scenario depicts "Classification", meaning that the predicted output will fall into one of two or more discrete categories -- "good" and "bad" in our case.


Article 1: Why Machine Learning? โ€“ Apurba Learns ML

#artificialintelligence

For those who aren't familiar with the show, it depicts an all-seeing AI that can predicts crime and other immoral acts before they even happen and passes on the information to the Government. One of the episodes of the show illustrates the origin of the Machine, where its creator -- Harold Finch -- is teaching it to distinguish between good and bad by showing it "examples". This is a perfect example of Supervised learning -- we have a data-set or example set with the right answers given. Afterwards, we expect the computer to predict things based on the data-set. Now, to be a bit more specific, the above scenario depicts "Classification", meaning that the predicted output will fall into one of two or more discrete categories -- "good" and "bad" in our case.


The Six Steps to Boosted Trees

#artificialintelligence

BigML is bringing Boosted Trees to our ever-growing suite of supervised learning techniques. Boosting is a variation on ensembles that aims to reduce bias, potentially leading to better performance than Bagging or Random Decision Forests. In our first blog post of this series of six posts about Boosted Trees, we saw a gentle introduction to Boosted Trees to get some context about what this new resource is and how it can help you solve your classification and regression problems. This post will take us further, into the detailed steps of how to use boosting with BigML. To learn from our data, we must first upload it.


Discriminate-and-Rectify Encoders: Learning from Image Transformation Sets

arXiv.org Machine Learning

The complexity of a learning task is increased by transformations in the input space that preserve class identity. Visual object recognition for example is affected by changes in viewpoint, scale, illumination or planar transformations. While drastically altering the visual appearance, these changes are orthogonal to recognition and should not be reflected in the representation or feature encoding used for learning. We introduce a framework for weakly supervised learning of image embeddings that are robust to transformations and selective to the class distribution, using sets of transforming examples (orbit sets), deep parametrizations and a novel orbit-based loss. The proposed loss combines a discriminative, contrastive part for orbits with a reconstruction error that learns to rectify orbit transformations. The learned embeddings are evaluated in distance metric-based tasks, such as one-shot classification under geometric transformations, as well as face verification and retrieval under more realistic visual variability. Our results suggest that orbit sets, suitably computed or observed, can be used for efficient, weakly-supervised learning of semantically relevant image embeddings.


The Crossover Process: Learnability and Data Protection from Inference Attacks

arXiv.org Machine Learning

It is usual to consider data protection and learnability as conflicting objectives. This is not always the case: we show how to jointly control inference --- seen as the attack --- and learnability by a noise-free process that mixes training examples, the Crossover Process (cp). One key point is that the cp~is typically able to alter joint distributions without touching on marginals, nor altering the sufficient statistic for the class. In other words, it saves (and sometimes improves) generalization for supervised learning, but can alter the relationship between covariates --- and therefore fool measures of nonlinear independence and causal inference into misleading ad-hoc conclusions. For example, a cp~can increase / decrease odds ratios, bring fairness or break fairness, tamper with disparate impact, strengthen, weaken or reverse causal directions, change observed statistical measures of dependence. For each of these, we quantify changes brought by a cp, as well as its statistical impact on generalization abilities via a new complexity measure that we call the Rademacher cp~complexity. Experiments on a dozen readily available domains validate the theory.