Goto

Collaborating Authors

 Information Retrieval


Finding Similar Music using Matrix Factorization

#artificialintelligence

In a previous post I wrote about how to build a'People Who Like This Also Like ...' feature for displaying lists of similar musicians. My goal was to show how simple Information Retrieval techniques can do a good job calculating lists of related artists. For instance, using BM25 distance on The Beatles shows the most similar artists being John Lennon and Paul McCartney. One interesting technique I didn't cover was using Matrix Factorization methods to reduce the dimensionality of the data before calculating the related artists. This kind of analysis can generate matches that are impossible to find with the techniques in my original post.


An Entity Resolution Primer

#artificialintelligence

My name is Jonathan Armoza and I am a data science intern at Neustar and a PhD candidate in English Literature at New York University. My work focuses on the development of computational text mining and visualization methods in the emerging field of digital humanities. The era of big data has created the need to develop techniques and mechanisms to not only handle large datasets, but to understand them as well. Much of this influx of information is about people, places, and things. Although some of that data is anonymized, there are a number of reasons we might want to understand how to associate those real world "entities" with their data points.


Will search engines fall to AI?

#artificialintelligence

Lately, there's been a rumble pretty much everywhere about artificial intelligence, digital personal assistants, the Internet of Things, wearables and apps for everything. I've even written about what the rise of digital assistants means to search. There are some who claim that these new technologies will render search obsolete, passed over for the convenience and joy of an always-available digital world. I think they are wrong. Instead of looking at a search engine as an ad platform, we need to remember what it actually does for people.


Europe is targeting Google under antitrust laws but missing the bigger picture

The Guardian

Google it today and you'll see that the European Commission has turned up the heat in its long-running probe into anti-competitive behaviour by the web's most popular search engine. EC competition chief, Margrethe Vestager, issued formal objections alleging that Google abuses its dominant position in the market of "general internet search". In particular, the EC claims that Google artificially boosts its own products in returning Google comparison shopping results in its service "Google Shopping", even if those products aren't the best or cheapest โ€“ the "most relevant", as the Commission puts it โ€“ for consumers. Since taking office in November 2014, Vestager has made the Google inquiry a top priority, signalling a willingness to consider court battles and hefty fines if Google and other digital giants don't fall into line with European competition law. In this, she has displayed a distinct shift from her predecessor, Joaquรญn Almunia, whose multiple attempts to achieve private settlement with Google fell apart a year ago, before descending into a political and economic boxing match.


Google Built AI Software with Human-Like English Skills

#artificialintelligence

Computers don't have the ability to understand English language, and to address this issue, this is where the Google's free AI software dubbed Parsey McParseface, based on SyntaxNet machine learning framework comes in. The search engine giant announced on Thursday that it will release a new piece of software that will understand written English. More interestingly, Google will release the AI software as open-source which means developers can freely use the software to train computer programs with ability to process natural language. Google says its latest innovation will be equipped with Natural Language Understanding (NLU) to parse written English sentences with up to 94 percent of accuracy. Google added that trained human linguists can only achieve up to 96 percent of accuracy. As an AI machine to achieve such milestone, it is a huge breakthrough in the field of artificial intelligence.


Google Removing Payday Loan Ads From Its Search Engine

International Business Times

Search giant Google announced Wednesday that it would cut payday loan providers from its advertising platforms, citing the potentially damaging effects to borrowers of short-term, high-interest cash loans. "Research has shown that these loans can result in unaffordable payment and high default rates for users so we will be updating our policies globally to reflect that," Google's head of global product policy, David Graf,f said in an announcement posted to the company's blog. The average payday loan borrower spends five months of the year in debt, paying more in fees than originally received, according to research compiled by the Pew Charitable Trusts. "Our hope is that fewer people will be exposed to misleading or harmful products," Graff said. The policy change, which follows a similar move by Facebook, won plaudits from advocacy groups concerned with the impact of payday loans on low-income borrowers.


Finding Similar Music using Matrix Factorization

@machinelearnbot

In a previous post I wrote about how to build a'People Who Like This Also Like ...' feature for displaying lists of similar musicians. My goal was to show how simple Information Retrieval techniques can do a good job calculating lists of related artists. For instance, using BM25 distance on The Beatles shows the most similar artists being John Lennon and Paul McCartney. One interesting technique I didn't cover was using Matrix Factorization methods to reduce the dimensionality of the data before calculating the related artists. This kind of analysis can generate matches that are impossible to find with the techniques in my original post.


The search engine for ARGUMENTS: Engineers plan to make a tool to help settle complex political online discussions

Daily Mail - Science & tech

These days, many an argument over trivial questions like the year a celebrity died or when an historical event happened can be settled with a quick search on Google. But search engines could soon be used to settle more complex arguments, involving serious political debate. A group of researchers in Germany are looking into ways to use search engines to shed light on the most complex political discussions. In just seconds, digital systems should be able to evaluate millions of documents, such as online discussions on controversial topics like the Transatlantic Trade and Investment Partnership (TTIP). The'Robust Argumentation Machines' project has been set up by Bielefeld University in Germany.


Drill Data with Apache Drill

@machinelearnbot

Apache Drill is a low-latency distributed query engine for large-scale datasets, including structured and semi-structured/nested data. Inspired by Google's Dremel, Drill is designed to scale to several thousands of nodes and query petabytes of data at interactive speeds that BI/Analytics environments require. Apache Drill includes a distributed execution environment, purpose built for large-scale data processing. At the core of Apache Drill is the "Drillbit" service which is responsible for accepting requests from the client, processing the queries, and returning results to the client. When a Drillbit runs on each data node in a cluster, Drill can maximize data locality during query execution without moving data over the network or between nodes.


SlashPixels: an ambitious image search engine for designers

#artificialintelligence

Google is so dominant in the search engine market at large that it becomes hard to launch anything that remotly looks like a search tool. A team of Russian developers decided to still give it a go and focus on a niche market: image search. The team's objective seems very ambitious, create an artificial intelligence based image search engine to help designers find inspiration or resources in an easier way. They promise that SlashPixels will understand each image that it indexes, thus giving it a big advantage when it comes to sort the pictures. Unfortunatly, all this doesn't exist yet, but you can support the team's IndieGogo campaign to help them build this new tool.