Goto

Collaborating Authors

 Oceania


Linear Dimensionality Reduction in Linear Time: Johnson-Lindenstrauss-type Guarantees for Random Subspace

arXiv.org Machine Learning

We consider the problem of efficient randomized dimensionality reduction with norm-preservation guarantees. Specifically we prove data-dependent Johnson-Lindenstrauss-type geometry preservation guarantees for Ho's random subspace method: When data satisfy a mild regularity condition -- the extent of which can be estimated by sampling from the data -- then random subspace approximately preserves the Euclidean geometry of the data with high probability. Our guarantees are of the same order as those for random projection, namely the required dimension for projection is logarithmic in the number of data points, but have a larger constant term in the bound which depends upon this regularity. A challenging situation is when the original data have a sparse representation, since this implies a very large projection dimension is required: We show how this situation can be improved for sparse binary data by applying an efficient `densifying' preprocessing, which neither changes the Euclidean geometry of the data nor requires an explicit matrix-matrix multiplication. We corroborate our theoretical findings with experiments on both dense and sparse high-dimensional datasets from several application domains.


Maximum Margin Principal Components

arXiv.org Machine Learning

Principal Component Analysis (PCA) is a very successful dimensionality reduction technique, widely used in predictive modeling. A key factor in its widespread use in this domain is the fact that the projection of a dataset onto its first $K$ principal components minimizes the sum of squared errors between the original data and the projected data over all possible rank $K$ projections. Thus, PCA provides optimal low-rank representations of data for least-squares linear regression under standard modeling assumptions. On the other hand, when the loss function for a prediction problem is not the least-squares error, PCA is typically a heuristic choice of dimensionality reduction -- in particular for classification problems under the zero-one loss. In this paper we target classification problems by proposing a straightforward alternative to PCA that aims to minimize the difference in margin distribution between the original and the projected data. Extensive experiments show that our simple approach typically outperforms PCA on any particular dataset, in terms of classification error, though this difference is not always statistically significant, and despite being a filter method is frequently competitive with Partial Least Squares (PLS) and Lasso on a wide range of datasets.


The Allen Institute of Artificial Intelligence (AI2) Joins Partnership on AI to Benefit People and Society :: ITbriefing.net ::

#artificialintelligence

We look forward to collaborating with other industry-leading Partnership on AI members to address the challenges and opportunities within the AI field including companies, nonprofits and institutions and with founding members Apple, Amazon, Facebook, Google / DeepMind, IBM and Microsoft; existing Partners AAAI, ACLU, OpenAI; and new Partners: AI Forum of New Zealand (AIFNZ), Allen Institute for Artificial Intelligence (AI2), Centre for Democracy & Tech (CDT), Centre for Internet and Society, India (CIS), Cogitai, Data & Society Research Institute (D&S), Digital Asia Hub, eBay, Electronic Frontier Foundation (EFF), Future of Humanity Institute (FHI), Future of Privacy Forum (FPF), Human Rights Watch (HRW), Intel, Leverhulme Centre for the Future of Intelligence (CFI), McKinsey & Company, SAP, Salesforce, Sony, UNICEF, Upturn, XPRIZE Foundation and Zalando.


Video Friday: Morphing Wheels, Soft Inflatable Robot, and Snipe Nano Quadrotor

IEEE Spectrum Robotics

Video Friday is your weekly selection of awesome robotics videos, collected by your Automaton bloggers. We'll also be posting a weekly calendar of upcoming robotics events for the next two months; here's what we have so far (send us your events!): Let us know if you have suggestions for next week, and enjoy today's videos. A very clever design for a small mobile robot by Draganfly Innovations: morphing wheels allow it to climb stairs with ease. The DraganScout by Draganfly Innovations Inc. is a unique ground-based robot with the ability to morph and adapt to different application or mission needs.


beamandrew/medical-data

#artificialintelligence

This is a curated list of medical data for machine learning. This list is provided for informational purposes only, please make sure you respect any and all usage restrictions for any of the data listed here. The National Library of Medicine presents MedPix Database of 53,000 medical images from 13,000 patients with annotations. These 1112 datasets are composed of structural and resting state functional MRI data along with an extensive array of phenotypic information. Also has clinical, genomic, and biomaker data. AMRG Cardiac Atlas The AMRG Cardiac MRI Atlas is a complete labelled MRI image set of a normal patient's heart acquired with the Auckland MRI Research Group's Siemens Avanto scanner.


Iteratively-Reweighted Least-Squares Fitting of Support Vector Machines: A Majorization--Minimization Algorithm Approach

arXiv.org Machine Learning

Support vector machines (SVMs) are an important tool in modern data analysis. Traditionally, support vector machines have been fitted via quadratic programming, either using purpose-built or off-the-shelf algorithms. We present an alternative approach to SVM fitting via the majorization--minimization (MM) paradigm. Algorithms that are derived via MM algorithm constructions can be shown to monotonically decrease their objectives at each iteration, as well as be globally convergent to stationary points. We demonstrate the construction of iteratively-reweighted least-squares (IRLS) algorithms, via the MM paradigm, for SVM risk minimization problems involving the hinge, least-square, squared-hinge, and logistic losses, and 1-norm, 2-norm, and elastic net penalizations. Successful implementations of our algorithms are presented via some numerical examples.


In a first, natural selection defeats a biocontrol insect

Science

Twenty years ago, Stephen Goldson thought he had beaten the Argentine stem weevil, an invasive insect that was devastating New Zealand's pastures. Goldson, an entomologist, had scoured the South American countryside and come up with an efficient weevil killer: a parasitoid wasp that at first killed up to 90% of the weevils. Now, the weevil has made a comeback and an examination of decades worth of data on its abundance over years has revealed that the weevil has outevolved its parasite, which produces asexually. Now, Goldson and his colleagues are studying weevil DNA to learn the secret of this comeback.


Using Deep Learning To Extract Knowledge From Job Descriptions

@machinelearnbot

At Search Party we are in the business of creating intelligent recruitment software. One of the problems we deal with is matching candidates and vacancies in order to create a recommendation engine. This usually requires parsing, interpreting and normalising messy, semi-/unstructured, textual data from rรฉsumรฉs and vacancies, which is where the following come in: conditional random fields, bag-of-words, TF-IDFs, WordNet, statistical analysis, but also a lot of manual work done by linguists and domain experts for the creation of synonym lists, skill taxonomies, job title hierarchies, knowledge bases or ontologies. While these concepts are valuable for the problem we try to solve, they also require a certain amount of manual feature engineering and human expertise. This expertise is certainly a factor that makes these techniques valuable, but the question remains whether more automated approaches can be used to extract knowledge about the job space to complement these more traditional approaches.


Why IT companies like Cognizant and Wipro are laying off employees

#artificialintelligence

The layoffs in India's IT companies, one of the largest private employers providing jobs to more than 37 lakh people, are worrying. Combined with the backlash the IT industry is facing in the US on the H1-B visa front, and in countries such as Australia, the future of the industry as a beacon of hope for young professionals is dimming, feel many. Add to this a widespread pessimism in the sector that much of its workforce will become redundant very soon (McKinsey in a report puts this number at half the workforce) and it is bad news for job-seekers and the government, which is staring at slack job creation across all segments. It also draws a grim picture of the economy's growth, which was being postured to be picking up smartly despite the demonetisation impact. Every business is going through huge transformation with the rampant use of technology that is making several traditional jobs obsolete.


Can an app use machine learning to inspire you to become more socially responsible? - Techly

#artificialintelligence

Acorns is the newest app capturing the imagination of the Australian market. Aiming to streamline the saving process while making it easier than ever to enter the investing sphere, Acorns is the hyped-up US micro-investing app which launched in Australia last year. A Techly Guest Post by venture capitalist, Omar Khan, had a look at the unique functions of the Acorns app. The app's investment options are broken down like so: "The app gives investors several options. The other, a voluntary contribution whenever I have some spare money I'd like to save.