Goto

Collaborating Authors

 Genre


Amazon.com: Mastering Apache Spark eBook: Mike Frampton: Kindle Store

@machinelearnbot

The book provides a super fast, short introduction to Spark in the first chapter and then jump straight into MLlib, Spark Streaming Spark SQL, GraphX, etc. in subsequent chapters. A huge positive for this book is that it not only talks about Spark itself, but also covers using Spark with other big data technologies like Hadoop, Kafka, Titan, Neo4j, HBase, Cassandra, H2O, etc. True to the name, sure the book covers more than simple introductory Spark topics, but it concentrates on breath than depth. There is decent coverage and enough code examples for each topic, but what it lacks is depth. There is no "best practices" or "performance" or "watch out for" type discussions or any type of advanced code. The MLlib chapter covers Naive Bayes, K-Means and Artificial Neural Networks (ANN).


8 Ways AI Will Profoundly Change City Life by 2030

#artificialintelligence

How will AI shape the average North American city by 2030? A panel of experts assembled as part of a century-long study into the impact of AI thinks its effects will be profound. The One Hundred Year Study on Artificial Intelligence is the brainchild of Eric Horvitz, a computer scientist, former president of the Association for the Advancement of Artificial Intelligence, and managing director of Microsoft Research's main Redmond lab. Every five years a panel of experts will assess the current state of AI and its future directions. The first panel, comprised of experts in AI, law, political science, policy, and economics, was launched last fall and decided to frame their report around the impact AI will have on the average American city.


[Discussion] I am following Andrew Ng's Coursera course. Is there an entry course to better follow it? • /r/MachineLearning

@machinelearnbot

I can't offer much in terms of other entry level recommendations, but I can recommend you learn to utilize the resource pages on the coursera course. The way the andrew NG course is set up is that you more or less try to have an idea of how these algorithms work at a conceptual level through the videos, then when you go to programming assignments, you can skip a lot of the prep work and focus on implementing the machine learning algorithms. Now those algorithms might be a little hard to follow at first, which is okay and expected, and that's where the lecture notes and/or wiki come in. From the wiki you can more or less translate the math formulas into code syntax and the assignments are more or less complete. The weeks build off each other so as you learn how to do one part, they do a little less prep work for you so you have to learn how to do another part, and so forth.


Stanford scientists develop novel brain-sensing technology that allows typing at 12 words per minute

#artificialintelligence

It does not take an infinite number of monkeys to type a passage of Shakespeare. Instead, it takes a single monkey equipped with brain-sensing technology - and a cheat sheet. That technology, developed by Stanford Bio-X scientists Krishna Shenoy, a professor of electrical engineering at Stanford, and postdoctoral fellow Paul Nuyujukian, directly reads brain signals to drive a cursor moving over a keyboard. In an experiment conducted with monkeys, the animals were able to transcribe passages from the New York Times and Hamlet at a rate of up to 12 words per minute. Earlier versions of the technology have already been tested successfully in people with paralysis, but the typing was slow and imprecise. This latest work tests improvements to the speed and accuracy of the technology that interprets brain signals and drives the cursor.


Reading: "Mining Large Streams of User Data for Personalized Recommendations"

#artificialintelligence

Data Scientists across Skyscanner have started meeting every fortnight to discuss research papers that tackle similar problems to those that we face within Skyscanner. The 2nd paper we read was: "Mining Large Streams of User Data for Personalized Recommendations" (hi Xavier!). Just like the last post, we're we're also writing up a brief, non-technical overview the problems/opportunities we discussed. Netflix famously announced a 1M prize in 2006, calling on researchers across the world to improve their movie recommender system by 10%. To create this competition, they had to make a critical decision: how could Netflix measure a 10% improvement in their system?


The World Series of Hacking--without humans

#artificialintelligence

LAS VEGAS--On a raised floor in a ballroom at the Paris Hotel, seven competitors stood silently. These combatants had fought since 9:00am, and nearly 4 million in prize money loomed over all the proceedings. Now some 10 hours later, their final rounds were being accompanied by all the play-by-play and color commentary you'd expect from an episode of American Ninja Warrior. Yet, no one in the competition showed signs of nerves. To observers, this all likely came across as odd--especially because the competitors weren't hackers, they were identical racks of high-performance computing and network gear.


Focused Model-Learning and Planning for Non-Gaussian Continuous State-Action Systems

arXiv.org Machine Learning

We introduce a framework for model learning and planning in stochastic domains with continuous state and action spaces and non-Gaussian transition models. It is efficient because (1) local models are estimated only when the planner requires them; (2) the planner focuses on the most relevant states to the current planning problem; and (3) the planner focuses on the most informative and/or high-value actions. Our theoretical analysis shows the validity and asymptotic optimality of the proposed approach. Empirically, we demonstrate the effectiveness of our algorithm on a simulated multi-modal pushing problem.


From Causes for Database Queries to Repairs and Model-Based Diagnosis and Back

arXiv.org Artificial Intelligence

In this work we establish and investigate connections between causes for query answers in databases, database repairs wrt. denial constraints, and consistency-based diagnosis. The first two are relatively new research areas in databases, and the third one is an established subject in knowledge representation. We show how to obtain database repairs from causes, and the other way around. Causality problems are formulated as diagnosis problems, and the diagnoses provide causes and their responsibilities. The vast body of research on database repairs can be applied to the newer problems of computing actual causes for query answers and their responsibilities. These connections, which are interesting per se, allow us, after a transition -inspired by consistency-based diagnosis- to computational problems on hitting sets and vertex covers in hypergraphs, to obtain several new algorithmic and complexity results for database causality.


Simpler PAC-Bayesian Bounds for Hostile Data

arXiv.org Machine Learning

Learning theory can be traced back to the late 60s and has attracted a great attention since. We refer to the monographs Devroye et al. (1996) and Vapnik (2000) for a survey. Most of the literature addresses the simplified case of i.i.d observations coupled with bounded loss functions. Many bounds on the excess risk holding with large probability were provided - these bounds are refered to as PAC learning bounds since Valiant (1984). In the late 90s, the PAC-Bayesian approach has been pioneered by Shawe-Taylor and Williamson (1997) and McAllester (1998, 1999). It consists in producing PAC bounds for a specific class of Bayesian-flavored estimators. Similarly to classical PAC results, most PAC-Bayesian bounds have been obtained with bounded loss functions (see Catoni, 2007, for some of the most accurate results). Note that Catoni (2004) provides bounds for unbouded loss, but still under very strong exponential moments assumptions. These assumptions were essentially not improved in the most recent works Guedj and Alquier (2013) and Bégin et al. (2016).


Boltzmann-Machine Learning of Prior Distributions of Binarized Natural Images

arXiv.org Machine Learning

Prior distributions of binarized natural images are learned by using a Boltzmann machine. According the results of this study, there emerges a structure with two sublattices in the interactions, and the nearest-neighbor and next-nearest-neighbor interactions correspondingly take two discriminative values, which reflects the individual characteristics of the three sets of pictures that we process. Meanwhile, in a longer spatial scale, a longer-range, although still rapidly decaying, ferromagnetic interaction commonly appears in all cases. The characteristic length scale of the interactions is universally up to approximately four lattice spacings $\xi \approx 4$. These results are derived by using the mean-field method, which effectively reduces the computational time required in a Boltzmann machine. An improved mean-field method called the Bethe approximation also gives the same results, as well as the Monte Carlo method does for small size images. These reinforce the validity of our analysis and findings. Relations to criticality, frustration, and simple-cell receptive fields are also discussed.