Goto

Collaborating Authors

 Information Retrieval


Exploiting Crowd-Based Labels for Domain Focused Information Retrieval

AAAI Conferences

Information search and retrieval from online sources or social forums is often performed with term based boolean queries. Such queries can produce low relevance documents in situations where the user is interested in retrieving in- formation related to a concept, or belonging to a specific domain. In this work an approach for concept-based infor- mation retrieval is presented, which exploits word and doc- ument distributions derived from topic modeling performed on data from online sources. Documents acquired from the Reddit and Stack Exchange online social forums are used for extracting concepts, and subsequently training and testing a detector that aids in identifying and retrieving documents associated with the concept of interest. The selection of training sets for our concept based detector is aided by pre-partitioning of documents by online users (or crowd) into concept focused sub-forums, such as sub-reddits. Topics derived from a sample of the overall document set are taken to represent concepts. These topics then form the basis for identifying sub-forums that have a strong correspondence with the concept of interest, and documents within are assigned (noisy) binary labels. The applicability of our approach is demonstrated by creating a domain focused detector for Cyber Security content from Reddit data. The cross utility of this detector is demonstrated by success- fully retrieving relevant Cyber Security documents from an alternate test online source: Stack Exchange. Document classification results of the proposed approach are compared favorably with classifications performed by human analysts.



China Investigates Search Engine Baidu After Student Dies Of Cancer

NPR Technology

Baidu, China's largest search engine, is under investigation after college student with a rare form of cancer said it promoted a fraudulent treatment center. Baidu, China's largest search engine, is under investigation after college student with a rare form of cancer said it promoted a fraudulent treatment center. Chinese health and Internet authorities have launched an investigation into Baidu, the country's largest search engine, following the death of a college student who accused Baidu of misleading him to a fraudulent cancer treatment. Experts believe the scandal will damage the credibility of Baidu's search results, and its long-term economic prospects. On Monday, news of the government investigation caused Baidu's stock to tumble by nearly 8% on the NASDAQ.


Two great ideas to create a much better search engine

@machinelearnbot

When you do a search for "career objectives" on Google India (www.google.in), the first result showing up is from a US-based job board specializing in data mining and analytical jobs. The Google link in question redirects to a page that does not even contain the string "career objective". In short, Google is pushing a US web site that has nothing to do with "career objectives" as the #1 web site for "career objectives" in India. In addition, Google totally failed to recognize that the web site in question is about analytics and data mining.


Deep Language Modeling for Question Answering using Keras

#artificialintelligence

Question answering has recieved more focus as large search engines have basically mastered general information retrieval and are starting to cover more edge cases. Question answering happens to be one of those edge cases, because it could involve a lot of syntatic nuance that doesn't get captured by standard information retrieval models, like LDA or LSI. Hypothetically, deep learning models would be better suited to this type of task because of their ability to capture higher-order syntax. Two papers, "Applying deep learning to answer selection: a study and an open task" (Feng et. Personally, I am a lot lazier than them, and I don't understand CNNs very well, so I would like to use an existing framework to build one of their models to see if I could get similar results. Keras is a really popular one that has support for everything we might need to put the model together. The Github repository for this project can be found here. See the instructions here on how to install Keras.


6 Best SEO Practices For Machine Learning - Shane Barker

#artificialintelligence

With the face of SEO constantly evolving, machine learning has become a huge concern for Internet marketers. An exciting post on the Moz blog about machine learning persuaded me to dig deeper into it. Eric Enge very clearly explained how machine learning works and ways Google may be using it. Google has been dominating the search engine world for decades, but this new concept may spark a complete overhaul to the Google spam-fighting algorithm updates. Internet marketers will have to adapt the best SEO practices according to this latest development.


How close are AI systems to human-level intelligence? The Allen AI challenge.

#artificialintelligence

With respect to artificial intelligence, some people are squarely in the "optimist" camp, believing that we are "nearly there" as far as producing human-level intelligence. Microsoft co-founder's Paul Allen has been somewhat more prudent: While we have learned a great deal about how to build individual AI systems that do seemingly intelligent things, our systems have always remained brittle--their performance boundaries are rigidly set by their internal assumptions and defining algorithms, they cannot generalize, and they frequently give nonsensical answers outside of their specific focus areas. So Allen does not believe that we will see human-level artificial intelligence in this century. But he nevertheless generously created a foundation aiming to develop such human-level intelligence, the Allen Institute for Artificial Intelligence Science. The Institute is lead by Oren Etzioni who obviously shares some of Allen's "pessimistic" views.


Search Engine Optimisation: Few Things to Know

@machinelearnbot

SEO or search engine optimisation is an internet marketing process to increase the placement of your website in search results found on search engines like Google and Bing. In order to make your website search engine friendly, SEO companies use some white-hat on-page techniques. In other words, SEO or search engine optimisation includes a set of rules, which are followed by blogs or website owners in order to optimise their websites for search engines. As a business owner one should know what the benefits of SEO services are. SEO is the best marketing strategy to secure your position in the Google algorithm.


You Could Look It Up by Jack Lynch review – search engines can't do everything

The Guardian

For some years now, the most satisfyingly passive-aggressive way of responding to a factual query on social media has been to reply with a link from the website "Let Me Google That For You". On opening the link, your pesterer sees an animation of their exact query being typed into the Google search field, the "I'm feeling lucky" box being clicked and a page showing what is almost certainly the answer to their question. It is a sadistically elaborate vehicle for a simple message: you are wasting both our time by asking a person something, when you could ask a search engine. But the search engine is hardly infallible. It is commonly assumed these days that all useful information is on the internet, but it isn't.


The Effects of Machine Learning on Rankings and SEO

#artificialintelligence

For a long time search engines relied on static ranking factors. Those webmasters and SEOs who knew what to pay attention for were able to reach the best positions on Google's SERPs. This has changed recently and will be changing in the future: The increasing usage of machine learning techniques leads to both dynamic ranking criteria and – as confusing as it may sound – a greater influence of human signals. Machine learning is nothing new. Its roots go back to the 50s of the last century.